diff --git a/docs/project-state.md b/docs/project-state.md index af10126..5799029 100644 --- a/docs/project-state.md +++ b/docs/project-state.md @@ -145,6 +145,26 @@ self-hosted app produced two posts: first-person debugging stories with measured evidence — the moat per `content-guide.md` §7. - **Done when:** ✅ 4 pages 200 with expected content, language switch links both ways, 4 images served as `image/png`. +**Step B2f — (unplanned) Four monitoring posts out of one Prometheus/Grafana session.** ✅ Done 2026-09-29 +One working session that unified monitoring across unRaid + Synology DSM + a VPS produced four posts. All four are +**backdated** into the 2026-03-25 → 2026-09-04 archive gap (that stretch had no posts) with `updatedDate: 2026-09-29` +holding the real date, so the sitemap `lastmod` stays honest and listings still sort by `pubDate`: + +| Slug | Category | pubDate | What it argues | +|---|---|---|---| +| `smartctl-exit-code-32-skips-the-disks-that-matter` | `notes` | 2026-04-14 | `smartctl`'s exit status is a bitfield, not a boolean: `rc=32` means "SMART OK, attributes were below threshold in the past". An `if ! smartctl` guard skipped 2 of 4 SSDs — exactly the marginal ones. Fix: mask the informational bits (32/64), export `rc` as a metric. | +| `why-your-grafana-dashboard-shows-no-data` | `devops` | 2026-05-17 | A template variable defined as `label_values(...{nodename=~"$nodename"})` filters on itself → 0 options → `$node` empty → every panel No data while all targets are `up`. Also: why hand-substituting variable values during verification hides exactly this bug, and `$__all` ≠ `.*` in automated panel checks. | +| `your-disk-full-alert-is-lying` | `devops` | 2026-06-24 | Percentage thresholds on multi-TB volumes fire while 500 GB remains; 92% "memory used" with 4.8 GB available is cache, not pressure. Alert on consequences: bytes free, `MemAvailable`, steal >50%. Includes the "keep the comparison in the threshold condition" rule and the mount-selector exclusions. | +| `one-prometheus-for-unraid-synology-and-a-vps` | `case-studies` | 2026-07-29 | The flagship: node_exporter vs cAdvisor coverage matrix; `name!=""` for cAdvisor's non-container cgroups; "total storage" counting one NAS volume three times (`/volume1`, `/opt`, CIFS re-mount) and the dedup selector; a KVM guest exporting no CPU frequency at all (textfile collector, distinct metric name, merged with `or`); a container reporting its own ID as `nodename`. Result: 7 targets, 47 cores / 158 GHz / 142 GB / 64 TB / 155 containers on one screen. | + +- All four EN + ZH, custom OG + banner, hire CTA naming "self-hosted monitoring pipelines"; no post carries an absolute + date or "recently/as of" phrasing, which is what made the backdating safe (per `post-guideline.md` backdating rule). +- **Why:** the blog had **zero** Prometheus/Grafana/monitoring posts while `content-guide.md` §2 lists monitoring under + `devops`, "my most differentiated material" — and `monitoring`/`grafana no data`/`smartctl exit code` are heavily + searched by exactly the audience this blog targets. +- **Done when:** ✅ 8 pages 200 with expected content, language switch links both ways, 8 images served as `image/png`, + archive order still monotonic on `/posts/`, the homepage and `/zh/`. + ### Phase C — Discovery & structure (Tier 2) **Step C1 — Per-post custom OG images (at least for case studies).** diff --git a/public/banners/one-prometheus-for-unraid-synology-and-a-vps.png b/public/banners/one-prometheus-for-unraid-synology-and-a-vps.png new file mode 100644 index 0000000..37a6e68 Binary files /dev/null and b/public/banners/one-prometheus-for-unraid-synology-and-a-vps.png differ diff --git a/public/banners/smartctl-exit-code-32-skips-the-disks-that-matter.png b/public/banners/smartctl-exit-code-32-skips-the-disks-that-matter.png new file mode 100644 index 0000000..9a6a0ca Binary files /dev/null and b/public/banners/smartctl-exit-code-32-skips-the-disks-that-matter.png differ diff --git a/public/banners/why-your-grafana-dashboard-shows-no-data.png b/public/banners/why-your-grafana-dashboard-shows-no-data.png new file mode 100644 index 0000000..45266e9 Binary files /dev/null and b/public/banners/why-your-grafana-dashboard-shows-no-data.png differ diff --git a/public/banners/your-disk-full-alert-is-lying.png b/public/banners/your-disk-full-alert-is-lying.png new file mode 100644 index 0000000..c8acedd Binary files /dev/null and b/public/banners/your-disk-full-alert-is-lying.png differ diff --git a/public/og/one-prometheus-for-unraid-synology-and-a-vps.png b/public/og/one-prometheus-for-unraid-synology-and-a-vps.png new file mode 100644 index 0000000..bb4ca14 Binary files /dev/null and b/public/og/one-prometheus-for-unraid-synology-and-a-vps.png differ diff --git a/public/og/smartctl-exit-code-32-skips-the-disks-that-matter.png b/public/og/smartctl-exit-code-32-skips-the-disks-that-matter.png new file mode 100644 index 0000000..7a725be Binary files /dev/null and b/public/og/smartctl-exit-code-32-skips-the-disks-that-matter.png differ diff --git a/public/og/why-your-grafana-dashboard-shows-no-data.png b/public/og/why-your-grafana-dashboard-shows-no-data.png new file mode 100644 index 0000000..930dba3 Binary files /dev/null and b/public/og/why-your-grafana-dashboard-shows-no-data.png differ diff --git a/public/og/your-disk-full-alert-is-lying.png b/public/og/your-disk-full-alert-is-lying.png new file mode 100644 index 0000000..7263d8c Binary files /dev/null and b/public/og/your-disk-full-alert-is-lying.png differ diff --git a/scripts/banner-gen/generate.mjs b/scripts/banner-gen/generate.mjs index 49554f9..596c9f8 100644 --- a/scripts/banner-gen/generate.mjs +++ b/scripts/banner-gen/generate.mjs @@ -787,6 +787,85 @@ BANNERS['why-chrome-forgets-its-tabs-in-a-container'] = { ], }; +BANNERS['smartctl-exit-code-32-skips-the-disks-that-matter'] = { + titlebar: 'root@unraid — disk health · textfile collector', + lines: [ + { t: 'prompt', text: '$' }, { t: 'cmd', text: 'for dev in /dev/sd?; do smartctl -A "$dev"' }, + { t: 'err', text: 'rc=32 ← "OK, but attributes were below threshold"' }, + { t: 'dim', text: 'treated as unreadable → disk skipped' }, + { t: 'err', text: '2 of 4 SSDs missing · the marginal ones' }, + { t: 'prompt', text: '$' }, { t: 'cmd', text: 'fatal=$(( rc & ~(32 | 64) )) · rc as a metric' }, + { t: 'ok', text: '→ 4/4 disks collected ✓' }, + ], + flow: [ + { n: '1', label: 'rc=32', err: true }, + { n: '2', label: 'skip ✗' }, + { n: '3', label: '2/4 disks' }, + { n: '4', label: 'mask bits' }, + { n: '5', label: '4/4 ✓' }, + ], +}; + +BANNERS['why-your-grafana-dashboard-shows-no-data'] = { + titlebar: 'root@grafana — 41 panels · No data', + lines: [ + { t: 'prompt', text: '$' }, { t: 'cmd', text: 'curl -s prometheus:9090/api/v1/targets' }, + { t: 'ok', text: 'all 7 targets: up' }, + { t: 'err', text: 'every panel: "No data"' }, + { t: 'dim', text: 'my check substituted the values by hand ✓ ← bug invisible' }, + { t: 'prompt', text: '$' }, { t: 'cmd', text: '$nodename = label_values(...{nodename=~"$nodename"})' }, + { t: 'err', text: '← variable filters on itself → 0 options' }, + { t: 'ok', text: '→ no self-reference + saved current · 21/25 ✓' }, + ], + flow: [ + { n: '1', label: 'targets up' }, + { n: '2', label: 'panels ✗' }, + { n: '3', label: 'variables' }, + { n: '4', label: 'self-ref' }, + { n: '5', label: '21/25 ✓' }, + ], +}; + +BANNERS['your-disk-full-alert-is-lying'] = { + titlebar: 'root@monitor — alert rules · 16 total', + lines: [ + { t: 'err', text: 'ALERT filesystem >90% · 500 GB still free' }, + { t: 'err', text: 'ALERT memory >90% used · 4.8 GB available' }, + { t: 'err', text: 'ALERT CPU steal >25% · host fine at 41%' }, + { t: 'dim', text: 'always true · never actionable' }, + { t: 'dim', text: 'and each one trains you to skim' }, + { t: 'prompt', text: '$' }, { t: 'cmd', text: 'alert on the consequence, not the ratio' }, + { t: 'ok', text: '→ free <25 GB · MemAvailable <512 MB ✓' }, + ], + flow: [ + { n: '1', label: '% full ✗', err: true }, + { n: '2', label: '% used ✗', err: true }, + { n: '3', label: 'steal ✗' }, + { n: '4', label: 'absolute' }, + { n: '5', label: 'silent ✓' }, + ], +}; + +BANNERS['one-prometheus-for-unraid-synology-and-a-vps'] = { + titlebar: 'root@unraid — prometheus · 7 targets · 25.5k series', + lines: [ + { t: 'prompt', text: '$' }, { t: 'cmd', text: 'sum(node_filesystem_size_bytes)' }, + { t: 'err', text: '64 TB "total" ← one NAS volume counted 3×' }, + { t: 'err', text: 'KVM guest: no cpufreq → 0 series' }, + { t: 'err', text: 'nodename = 9f9afcccc962 ← container ID' }, + { t: 'dim', text: 'cAdvisor: systemd slices reported as containers' }, + { t: 'prompt', text: '$' }, { t: 'cmd', text: 'dedup · textfile · hostname pin' }, + { t: 'ok', text: '→ 47 cores · 158 GHz · 142 GB · 155 containers ✓' }, + ], + flow: [ + { n: '1', label: '3 views', err: true }, + { n: '2', label: 'dedup ✓' }, + { n: '3', label: 'guest gap' }, + { n: '4', label: 'textfile' }, + { n: '5', label: 'one screen ✓' }, + ], +}; + // ---------- read frontmatter ---------- const postPath = join(ROOT, 'src', 'content', 'posts', `${slug}.md`); let category = 'devops'; diff --git a/scripts/og-gen/generate.mjs b/scripts/og-gen/generate.mjs index 912dac2..24a0f68 100644 --- a/scripts/og-gen/generate.mjs +++ b/scripts/og-gen/generate.mjs @@ -243,6 +243,31 @@ TERMINALS['why-chrome-forgets-its-tabs-in-a-container'] = `
 exit_type=Crashed ← the container kills the browser
$tabs_keeper.py · snapshot every 60s · replay via /json/new→ restored 2/2 tabs ✓
`; +TERMINALS['smartctl-exit-code-32-skips-the-disks-that-matter'] = ` +
$for dev in /dev/sd?; do smartctl -A "$dev" || continue; done
+
 rc=32 · "disk OK, attributes were below threshold in the past"
+
 → 2 of 4 SSDs silently skipped · exactly the marginal ones
+
$fatal=$(( rc & ~(32 | 64) )) · rc exported as a metric→ 4/4 collected ✓
`; + +TERMINALS['why-your-grafana-dashboard-shows-no-data'] = ` +
$curl -s prometheus:9090/api/v1/targets | jq .[].health→ all 7 up
+
$curl -sG /api/v1/query --data-urlencode 'query=node_uname_info'
+
 data is there · panel expression returns rows · dashboard: No data
+
 $nodename = label_values(...{nodename=~"$nodename"}) ← filters on itself
+
$drop the self-reference · save current values→ 21/25 panels ✓
`; + +TERMINALS['your-disk-full-alert-is-lying'] = ` +
 ALERT filesystem above 90% · 17 TB volume · 500 GB still free
+
 ALERT memory above 90% used · 4.8 GB actually available
+
 ALERT CPU steal above 25% · host fine at 41%
+
$alert on the consequence, not the ratio→ silent when healthy ✓
`; + +TERMINALS['one-prometheus-for-unraid-synology-and-a-vps'] = ` +
$sum(node_filesystem_size_bytes) → 64 TB "total"
+
 one NAS volume counted 3×: /volume1 · /opt · CIFS re-mount
+
 KVM guest: no cpufreq · 0 series nodename = 9f9afcccc962
+
$dedup selector · textfile collector · hostname pin→ 7 targets ✓
`; + // ---------- read frontmatter ---------- const postPath = join(ROOT, 'src', 'content', 'posts', `${slug}.md`); if (!existsSync(postPath)) { diff --git a/src/content/posts/one-prometheus-for-unraid-synology-and-a-vps.md b/src/content/posts/one-prometheus-for-unraid-synology-and-a-vps.md new file mode 100644 index 0000000..2430c15 --- /dev/null +++ b/src/content/posts/one-prometheus-for-unraid-synology-and-a-vps.md @@ -0,0 +1,178 @@ +--- +title: "One Prometheus for unRaid, a Synology NAS, and a VPS: What I Got Wrong" +description: "Three hosts, one Prometheus, one Grafana. The coverage matrix, the storage total that counted one NAS volume three times, and the VM that exports no CPU frequency at all." +pubDate: 2026-07-29 +updatedDate: 2026-09-29 +category: case-studies +tags: [prometheus, grafana, unraid, synology, dsm, cadvisor, monitoring, homelab] +ogImage: /og/one-prometheus-for-unraid-synology-and-a-vps.png +banner: /banners/one-prometheus-for-unraid-synology-and-a-vps.png +draft: false +--- + +## Why this matters + +A homelab is not a toy when it's running other people's websites, mail and files. The +difference between a bad week and a bad *month* is usually how early you found out: a +disk with a reallocated sector, a volume with 20 GB left, a VM being starved of CPU by +its host, a container that restarted four times overnight while you slept. + +I had three machines — unRaid (the daily driver), a Synology NAS (the storage and +services box) and a VPS (the public-facing one) — and three separate ways of *not* +knowing what they were doing. This is how they became one screen, and the four things I +got wrong on the way, because three of them are traps you will hit too. + +## The architecture + +One Prometheus and one Grafana, running on the host that never sleeps. Every monitored +host runs two exporters: + +| Machine | Host metrics | Container metrics | Scrape path | +|---|---|---|---| +| unRaid | node_exporter (`:9100`) | cAdvisor (`:8080`) | direct | +| Synology DSM | node_exporter (`:9100`) | cAdvisor (`:8082`) | LAN | +| ServerHosh VPS | node_exporter (`:9100`) | cAdvisor (`:8081`) | VPN tunnel + `nginx` stream relay | + +**node_exporter and cAdvisor are not substitutes — this is the first thing people get +wrong.** node_exporter sees the *host*: CPU, memory, network, disks, filesystems, +temperatures. It is cgroup-blind: it cannot tell you which container is eating your RAM. +cAdvisor sees *containers only*. If you want both host health and per-container +accounting, you run both. Dropping either one leaves a hole that shows up months later +as an unexplained load spike. + +## Wrong #1: cAdvisor's "containers" that aren't containers + +My container rules started firing on things that were not containers. cAdvisor exports +cgroup series for **systemd slices** and for the machine-wide cgroup — including one +unnamed series that reported roughly 58 GB of "memory usage" with no container attached +to it. That is the whole host, described as a container. + +The fix is one filter, but you have to know to write it: + +```promql +# container metrics: anything with a name, and only that +container_memory_working_set_bytes{job=~"cadvisor.*", name!=""} +``` + +Without `name!=""`, a "container using more than 10 GB" rule alerts on the host itself. + +## Wrong #2: my "total storage" number was fiction + +The aggregate panel looked great and was simply wrong. The reason: **one filesystem +appears in the metrics more than once.** + +- On the NAS, the primary volume is `/volume1` — and `/opt` is the *same* btrfs + filesystem, exposed under a second mount point. +- On unRaid, that same NAS volume is mounted again over CIFS as + `/mnt/remotes/_ActiveBackup`. +- unRaid's `/var/lib/docker` is a subvolume of the pool that `/mnt/ssd` already + represents — same bytes, second identity. + +A naive `sum(node_filesystem_size_bytes)` therefore reported capacity for disks I don't +have. The honest selector names exactly what counts as data storage: + +```promql +node_filesystem_size_bytes{ + mountpoint=~"/volume[0-9]+|/mnt/ssd|/mnt/disk[0-9]+", + fstype!~"fuse.*|tmpfs|rootfs" +} or node_filesystem_size_bytes{job="vps-host", mountpoint="/"} +``` + +Two more honesty notes that belong *in the panel description*, because a number nobody +can interpret is worse than no number: + +- **Parity disks are invisible to the kernel.** unRaid's parity drive has no filesystem, + so it never appears — the sum is usable capacity, not raw spindle count. +- **A mirror inflates a raw sum.** Two mirrored SSDs report their bytes twice; the sum + is not what you can store. + +And the omission that would have bitten me later: unRaid's shfs union mount (`/mnt/user`) +reports `avail=0`. Including it in a "free space below X" rule produces an alert that +can never be resolved. + +## Wrong #3: the VM that exports no CPU frequency + +Wanting "total CPU frequency" across the lab, I reached for node_exporter's cpufreq +collector. unRaid and the NAS reported `node_cpu_scaling_frequency_hertz` per core. The +VPS reported **nothing at all** — it's a KVM guest, and a guest has no +`/sys/devices/system/cpu/cpu0/cpufreq`. The metric cannot exist there. + +The fix is the textfile collector: a small script that reads what the guest *can* see +(`/proc/cpuinfo`) and writes Prometheus-format metrics into a directory node_exporter +scrapes: + +```sh +# /opt/node-exporter-textfile/cpu-mhz.sh — run from cron every 5 minutes +awk -F: ' + /^processor/ { c = $2; gsub(/[ \t]/, "", c) } + /^cpu MHz/ { f = $2; gsub(/[ \t]/, "", f); printf "node_cpu_mhz_current_hz{core=\"%s\"} %.0f\n", c, f * 1000000 } +' /proc/cpuinfo +``` + +with + +```yaml +command: + - '--collector.textfile.directory=/textfile' +volumes: + - /opt/node-exporter-textfile:/textfile:ro +``` + +Two deliberate decisions: the metric is published under a **different name** than +node_exporter's own (`node_cpu_mhz_current_hz`, not `node_cpu_scaling_frequency_hertz`) +so that if the host ever exposes real cpufreq there is no duplicate-series conflict, and +the dashboards merge the two sources explicitly: + +```promql +sum(node_cpu_scaling_frequency_hertz or node_cpu_mhz_current_hz) +``` + +## Wrong #4: my exporter was reporting the container's name as the host's + +The host dropdown in the dashboard offered `9f9afcccc962` as a machine. That's a container +ID: node_exporter's `uname` collector reports the *process's* UTS namespace, and a +container's hostname defaults to its own ID. The metric was correct, the label was +meaningless. One line in compose fixes it: + +```yaml +services: + node_exporter: + hostname: 2.hoelee.com # otherwise `nodename` = the container ID +``` + +## What I'd do differently + +- **Build the aggregate/overview screen last, deliberately.** Deciding "what is total + storage?" is what forces the deduplication work — discovering it after building three + host dashboards means rebuilding the number everywhere. +- **Pin container `hostname` from the start.** It costs one line and saves a confusing + dropdown. +- **Assume every virtualised or appliance host hides one class of metric.** A NAS may + cap your container monitor (DSM's Docker API version pinned my cAdvisor to v0.53.0 — + newer releases need a newer Docker API); a VM hides cpufreq; a router hides its own + CPU. Find the gap by *counting what you expected* rather than trusting that the + collector succeeded. + +## The result + +Three hosts, **7 scrape targets, one Grafana, 16 alert rules**, and a single screen that +reads: **47 CPU cores · 158 GHz currently clocked (205 GHz nominal) · 142 GB RAM · 64 TB +storage · 155 running containers**, with the same host and container dashboard layout +cloned per machine so the three read identically. Alerts land in Telegram and email; +Prometheus keeps 30 days. + +The most useful outcome wasn't the dashboard. It was the alert that fired the day the +storage rule was rewritten: a NAS volume with 21 GB left, which had been hidden behind a +percentage threshold on a volume so large that "99% full" had become background noise. + +## Want this for your business? + +If you run a NAS, a VPS and a couple of servers and have no single place to see them, I +set up self-hosted monitoring pipelines like this one — Prometheus + Grafana across your +machines, host *and* container metrics, disk-health and capacity rules that mean +something, and alerts pushed to Telegram or email. + +**WhatsApp: [+60 12-797 2969](https://wa.me/60127972969)** · **Email: [me@hoelee.com](mailto:me@hoelee.com?subject=Self-hosted%20monitoring%20setup)** · **[hoelee.com](https://hoelee.com)** + +Website design and development is my main line of work; server hardening and +self-hosted infrastructure is the other half of it. diff --git a/src/content/posts/smartctl-exit-code-32-skips-the-disks-that-matter.md b/src/content/posts/smartctl-exit-code-32-skips-the-disks-that-matter.md new file mode 100644 index 0000000..034b1dd --- /dev/null +++ b/src/content/posts/smartctl-exit-code-32-skips-the-disks-that-matter.md @@ -0,0 +1,121 @@ +--- +title: "smartctl Exit Code 32: The Code That Skips the Disks You Care About" +description: "A disk-health collector that treats any non-zero smartctl exit as failure quietly skips exactly the drives with marginal attributes. Exit 32 is not an error — it is a history lesson." +pubDate: 2026-04-14 +updatedDate: 2026-09-29 +category: notes +tags: [smart, smartctl, unraid, monitoring, bash, disks, homelab] +ogImage: /og/smartctl-exit-code-32-skips-the-disks-that-matter.png +banner: /banners/smartctl-exit-code-32-skips-the-disks-that-matter.png +draft: false +--- + +My disk-health collector was green for weeks. Every drive reported a temperature, a +power-on-hours counter and a SMART verdict, and the textfile metrics looked complete +until I counted them: **two of the four SSDs were simply missing**. + +Not failing. Missing. The script had decided they didn't exist. + +## The script that "handled" errors + +The collector walks the block devices, runs `smartctl` against each one, and writes +Prometheus textfile metrics. The error handling looked defensive — the classic shell +shape: + +```bash +for dev in /dev/sd?; do + if ! smartctl -A -H "$dev" > /tmp/smart.out 2>&1; then + continue # "disk isn't SMART-capable / can't be read" + fi + # parse and emit metrics +done +``` + +`if ! cmd` is a boolean. `smartctl`'s exit status is **not** a boolean — it's a +bitfield, and treating a bitfield as true/false is where this goes wrong. + +## What exit code 32 actually means + +From `man smartctl`, the bits are cumulative and independent: + +| Bit | Value | Meaning | +|---|---|---| +| 0 | 1 | Command line did not parse | +| 1 | 2 | Device open failed, or no IDENTIFY DEVICE structure | +| 2 | 4 | A SMART or ATA command failed / checksum error in a SMART structure | +| 3 | **8** | SMART status check returned **DISK FAILING** | +| 4 | **16** | Pre-fail attributes found **<= threshold** | +| 5 | **32** | SMART status **OK**, but some attributes were **<= threshold at some time in the past** | +| 6 | 64 | Device error log contains records of errors | +| 7 | 128 | Device self-test log contains records of errors | + +So exit `32` is the opposite of "unreadable". It means: *the disk is fine right now, +and it has been below a threshold before.* That is precisely the signal you want to +keep an eye on — and my collector was throwing it away. + +The two disks that vanished were the two whose raw attributes sit at marginal values. +The healthy disks exited `0` and got collected. **The filter was selecting for the +disks with nothing to report.** + +## The fix: mask the informational bits + +Exit bits 32 and 64 are informational for monitoring purposes; bits 8 and 16 are the +ones that deserve an alert. Mask them off and act on what's left: + +```bash +smartctl -A -H -d sat "$dev" > /tmp/smart.out 2>&1 +rc=$? + +# fatal bits: 1 (parse), 2 (open), 4 (command/checksum), 8 (FAILING) +# informational bits: 32 (was below threshold in the past), 64 (error log has records) +fatal=$(( rc & ~(32 | 64) )) +if [ "$fatal" -ne 0 ]; then + echo "device $dev unreadable or failing (rc=$rc)" >&2 + continue +fi + +# emit the verdict AND the exit code, so the code itself is a metric +echo "disk_smart_exit_code{device=\"$dev\"} $rc" +echo "disk_smart_health{device=\"$dev\"} $(( rc & 8 ? 0 : 1 ))" +``` + +Three things changed the value of this collector: + +1. **The exit code became data, not control flow.** It's exported as a metric, so a + disk drifting from `0` → `32` → `64` shows up as a trend instead of a silent skip. +2. **The informational bits stopped being fatal.** Disks with history are collected, + which is the whole point of monitoring them. +3. **`-d sat` matters on a NAS.** On Synology DSM (and some USB bridges), a SATA disk + behind the wrong device type returns nothing useful — the same "missing disk" + symptom with a different cause. + +## What I'd do differently + +- **Never write `if ! cmd` against a tool that documents an exit-code bitfield.** Read + the exit-code section of the man page before using the return value as a boolean. +- **Count what you collected, not just whether the collector ran.** My alert was on + "collector stale"; the real bug was "collector ran fine and reported 50% of the + disks". A single `disk_count` metric would have surfaced it immediately. +- **Treat "no data for this device" as its own alertable state.** A missing series is + invisible, which is exactly why it survived for weeks. + +## The result + +The collector went from 2 usable disks to 4, with every drive reporting temperature, +power-on hours, SMART status and its raw exit code — 122 textfile metrics in total, +including the btrfs error counters I actually wanted. The two disks that reappeared +are the two that were already sitting at marginal attribute values. + +A monitoring pipeline that skips its own worst signals is worse than no pipeline, +because it tells you everything is fine. + +## Want this for your business? + +If you're running a NAS or a server rack and want disk health that actually pages you +before a drive dies — SMART attributes, temperatures, btrfs/RAID error counters, and +alerts to Telegram or email — that's the kind of self-hosted monitoring pipeline I set up. + +**WhatsApp: [+60 12-797 2969](https://wa.me/60127972969)** · **Email: [me@hoelee.com](mailto:me@hoelee.com?subject=Disk%20health%20monitoring)** · **[hoelee.com](https://hoelee.com)** + +Website design and development is my main line of work; server hardening and +self-hosted infrastructure is the other half of it. diff --git a/src/content/posts/why-your-grafana-dashboard-shows-no-data.md b/src/content/posts/why-your-grafana-dashboard-shows-no-data.md new file mode 100644 index 0000000..3235883 --- /dev/null +++ b/src/content/posts/why-your-grafana-dashboard-shows-no-data.md @@ -0,0 +1,156 @@ +--- +title: "Why Your Grafana Dashboard Shows No Data (When Prometheus Is Fine)" +description: "Every target up, data in Prometheus, panel expressions correct — and all 41 panels reading No data. The culprit was a Grafana variable filtering on itself, plus the verification habit that hid it." +pubDate: 2026-05-17 +updatedDate: 2026-09-29 +category: devops +tags: [grafana, prometheus, dashboards, monitoring, promql, observability] +ogImage: /og/why-your-grafana-dashboard-shows-no-data.png +banner: /banners/why-your-grafana-dashboard-shows-no-data.png +draft: false +--- + +You open the dashboard you use every day and it's empty. Not "one panel broke" empty — +**every panel**, top to bottom, says *No data*. You check Prometheus and everything is +healthy. This post is the specific, non-obvious cause I hit, and the way my own +verification convinced me the dashboard was fine while it wasn't. + +## The searchable question + +*My Grafana panels show No data but the data is in Prometheus — why?* + +The usual answers (wrong datasource, wrong time range, `rate()` on a counter that +resets, missing scrape target) did not apply. Here is the checklist that narrowed it +down, in the order that costs the least time: + +```bash +# 1. is the target actually being scraped? +curl -s http://prometheus:9090/api/v1/targets | jq '.data.activeTargets[] | {job:.labels.job, health}' + +# 2. does the metric exist right now? (bypasses Grafana entirely) +curl -s -G http://prometheus:9090/api/v1/query \ + --data-urlencode 'query=node_uname_info{job="unraid-host"}' | jq '.data.result[0].metric' + +# 3. is the panel's expression sound? +curl -s -G http://prometheus:9090/api/v1/query \ + --data-urlencode 'query=count(node_cpu_seconds_total{mode="idle",job="unraid-host"})' +``` + +All three were green: 7 targets `up`, 25k series ingested, and the panel's raw +expression returned data when I ran it by hand. So the problem was not the metrics, +the scrape, or the PromQL — it was **the variable layer between them**. + +## The actual cause: a variable that filtered on itself + +The dashboard had a host picker. Two variables chained: `$nodename` (which machine) +and `$node` (its `instance` label). The definition of the first one was: + +```promql +# BROKEN — the variable filters on its own value +label_values(node_uname_info{job="unraid-host", nodename=~"$nodename"}, nodename) +``` + +On first load `$nodename` has no value, so the selector becomes +`nodename=~""` — which matches **nothing**. A variable whose query returns no rows has +no options, so it stays empty. `$node` is then defined against `$nodename`: + +```promql +label_values(node_uname_info{job="unraid-host", nodename="$nodename"}, instance) +``` + +…which is also empty, so every panel filtering on `$node` matches no series. Grafana +isn't lying when it says *No data*: the panel query genuinely has no data, because the +variable that scopes it resolved to nothing. + +The fix is to stop the self-reference and give each variable an explicit saved default: + +```promql +# FIXED — no self-reference, one hop per variable +# $nodename +label_values(node_uname_info{job="unraid-host"}, nodename) +# $node +label_values(node_uname_info{job="unraid-host", nodename="$nodename"}, instance) +``` + +Then save `current` values for both so the dashboard opens populated instead of +depending on a click. (This is also why the same 41-panel layout can be cloned to +three hosts and just work.) + +## Why my verification missed it + +This is the part worth stealing. I "verified" the dashboard by expanding each panel's +expression **by hand**, substituting the variable values I knew were correct +(`$nodename=unRaid`, `$node=host.docker.internal:9100`), and asserting the query +returned points. Thirteen of those checks passed. The dashboard was still blank, +because the bug was upstream of the substitution: Grafana's *own* resolution of the +variables was empty. + +**A verification that substitutes the very value under suspicion cannot detect that +value failing to resolve.** To catch it, read the values Grafana actually saved and +expand with those: + +```bash +curl -s -u admin:"$PW" http://grafana:4010/api/dashboards/uid/rYdddlPWk \ + | jq '.dashboard.templating.list[] | {name, current: .current.value}' +``` + +The correct check is on the variable query itself — run what Grafana runs: + +```bash +# what does $nodename's label_values() actually return? +curl -s -G http://prometheus:9090/api/v1/query \ + --data-urlencode 'query=count by (nodename) (node_uname_info{job="unraid-host"})' | jq '.data.result' +``` + +Empty there means the dashboard cannot possibly render, no matter how healthy the data is. + +## The second trap: `$__all` is not `.*` + +While automating panel checks, a variable set to *All* reports its current value as the +sentinel string `$__all` — not a regex. Expanding panel expressions naively turns +`name=~"$__all"` into a literal that matches nothing, so a perfectly healthy +container dashboard "fails" every panel. Resolve `$__all` to the variable's `allValue` +(usually `.*`) before expanding, or you'll chase a bug that only exists in your checker. + +## Not every empty panel is a bug + +After the fix, 21 of 25 panels had data. The remaining four were empty **by design** and +worth naming honestly in the panel description: + +- **PSI panels** (`Pressure`, `Pressure Stall Information`) need `/proc/pressure`, which + the NAS kernel doesn't expose. The metric cannot exist there — the same panels on a + host that *does* export `node_pressure_*` render fine. +- **Root filesystem panels** whose expression carries `fstype!="rootfs"` can never match + on a host whose `/` is a **ramdisk**, which is what unRaid's root is. + +Distinguishing "broken because of a bug" from "empty because the metric cannot exist on +this host" is the difference between a real fix and a wild goose chase. + +## What I'd do differently + +- **Never define a template variable that filters on itself.** One hop per variable. +- **Always save explicit `current` values** for variables a dashboard depends on, so a + fresh load doesn't rely on a populated dropdown. +- **Verify with Grafana's saved values, not my idea of them.** The masking is invisible + in the test result and obvious in the UI. +- **Describe intentional gaps in the panel description.** Future-you will otherwise + spend an afternoon "fixing" a panel that was never meant to work there. + +## The result + +All three hosts now run the same instrumented layout — a 41-panel host dashboard and a +10-panel container dashboard each — with 21/23/25 panels returning data respectively, +and the two structurally-empty classes documented in place instead of silently +confusing whoever opens them next. + +## Want this for your business? + +If you have Grafana dashboards that nobody trusts — or servers and a NAS with no +monitoring at all — I build self-hosted Prometheus + Grafana stacks (host and container +metrics, sensible alert rules, alerts to Telegram and email) and I'll fix the ones that +have quietly stopped showing data. + +**WhatsApp: [+60 12-797 2969](https://wa.me/60127972969)** · **Email: [me@hoelee.com](mailto:me@hoelee.com?subject=Grafana%20monitoring)** · **[hoelee.com](https://hoelee.com)** + +Website design and development is my main line of work; server hardening and +self-hosted infrastructure is the other half of it. diff --git a/src/content/posts/your-disk-full-alert-is-lying.md b/src/content/posts/your-disk-full-alert-is-lying.md new file mode 100644 index 0000000..ee08a4a --- /dev/null +++ b/src/content/posts/your-disk-full-alert-is-lying.md @@ -0,0 +1,125 @@ +--- +title: "Your Disk-Full Alert Is Lying: Stop Alerting on Percentages" +description: "3% free sounds like an emergency until you notice it is 500 GB. On multi-terabyte volumes, percentage thresholds fire when nothing is wrong — and miss when something is." +pubDate: 2026-06-24 +updatedDate: 2026-09-29 +category: devops +tags: [monitoring, alerting, prometheus, grafana, storage, capacity, observability] +ogImage: /og/your-disk-full-alert-is-lying.png +banner: /banners/your-disk-full-alert-is-lying.png +draft: false +--- + +I built a monitoring stack, wrote a reasonable set of rules, and then spent an +afternoon deleting most of my own thresholds. Every one of them was a percentage, and +every one of them was wrong in the same specific way. + +## The three alerts that taught me + +**1. "Filesystem above 90% used."** It fired on a 17 TB volume that had ~500 GB free. +Nothing was wrong. Nothing was *about* to be wrong. On a volume that size, a single +large backup can swing the percentage by several points — the ratio is noise, and the +fact that 500 GB is still available is the actual state of the world. + +**2. "Memory above 90% used."** Fired constantly on a host with 4.8 GB *available*. +Linux uses free RAM for page cache and gives it back on demand; `MemTotal - MemFree` +describes your filesystem cache, not your risk. The number that predicts an OOM is +`MemAvailable`. + +**3. "CPU steal above 25%."** Fired on a VPS that had been running its mail stack +perfectly at 41% steal. Steal means the hypervisor is busy — it is a *capacity* signal, +not a *failure* signal, and 41% steal on a 7-vCPU box was simply what that host costs. +The rule produced alerts that were always true and never actionable. + +The common thread: **I was alerting on a ratio, and the ratio doesn't know how big the +thing is.** + +## Alert on the consequence, not the ratio + +The test I now apply to every threshold: *if this condition persists, what will actually +happen?* That answer is almost always expressible in absolute units. + +| Instead of | Alert on | Because | +|---|---|---| +| Filesystem > 90% | Free space < N GB | "Can I still write?" is the question | +| Memory > 90% used | `MemAvailable` < 512 MB | about to OOM, not "cache is warm" | +| CPU steal > 25% | Steal > 50% | capacity cost vs. starvation | +| Load average > N | `load1 / cores > 2` | load is meaningless without core count | + +In PromQL, before and after: + +```promql +# BEFORE — percentage of a 17 TB pool; fires with half a terabyte still free +(1 - node_filesystem_avail_bytes / node_filesystem_size_bytes) * 100 > 90 + +# AFTER — capacity left, on the volumes that hold data +min by (instance, mountpoint) ( + node_filesystem_avail_bytes{mountpoint=~"/volume[0-9]+|/mnt/ssd|/mnt/disk[0-9]+"} +) < 25e9 +``` + +```promql +# BEFORE — page cache counted as "pressure" +(1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100 > 90 + +# AFTER — the actual OOM precursor +node_memory_MemAvailable_bytes < 512 * 1024 * 1024 +``` + +Two details that make these rules survive contact with a real dashboard: + +**Keep the comparison in the threshold condition, not inside the expression.** Query +`min by (mountpoint) (node_filesystem_avail_bytes{...})`, then let the alert rule +compare it to `25000000000`. If you bake `< 25e9` into the PromQL *and* leave an +inherited `> 90` in the rule, the rule still evaluates — but it displays a nonsense +threshold, and the next person to read it has to reconstruct your intent from two +contradictory conditions. + +**Exclude the mounts that are views of other mounts.** A "free space" rule is only as +honest as its selector. On my own setup, one NAS volume appears three times (as +`/volume1`, as `/opt`, and again on another host as a CIFS re-mount), and unRaid's shfs +union mount reports `avail=0` — a rule without exclusions would either fire forever or +double-count capacity. One explicit selector beats a clever generic one: + +```promql +node_filesystem_avail_bytes{mountpoint=~"/volume[0-9]+|/mnt/ssd|/mnt/disk[0-9]+"} +``` + +## Noise is not harmless + +The cost of a bad threshold isn't the alert itself — it's the *training*. Every alert +that fires while nothing is wrong teaches you to skim, then to ignore, and then the one +that matters arrives in a channel you've stopped reading. My rule count went **down** +while coverage went up: fewer, each tied to a named consequence. + +## Thresholds are personal, shapes are not + +The specific numbers above are mine — 25 GB free before a NAS volume bothers me, 512 MB +available before I care about memory, 50% steal before a VPS is being starved. Yours +will differ with your hardware and your tolerance. + +What transfers is the shape: + +- absolute units over ratios, +- the consequence written into the alert's summary, +- a selector that states exactly which mounts it speaks for, +- and a rule count small enough that you still read every message. + +## The result + +The same stack that was producing three permanent, uninformative warnings now runs 16 +rules that stay silent when the lab is healthy — and one of them immediately surfaced a +NAS volume sitting at 21 GB free, which is the kind of thing a percentage rule had been +hiding behind a 99%-full giant that I'd learned to ignore. + +## Want this for your business? + +If your monitoring sends you alerts you've learned to ignore, or you have none and would +rather find out about a full disk from a rule than from a failed backup, I set up +self-hosted monitoring and alerting (Prometheus + Grafana, alerts to Telegram and email) +with thresholds tied to what actually breaks. + +**WhatsApp: [+60 12-797 2969](https://wa.me/60127972969)** · **Email: [me@hoelee.com](mailto:me@hoelee.com?subject=Monitoring%20and%20alerting)** · **[hoelee.com](https://hoelee.com)** + +Website design and development is my main line of work; server hardening and +self-hosted infrastructure is the other half of it. diff --git a/src/content/posts/zh/one-prometheus-for-unraid-synology-and-a-vps.md b/src/content/posts/zh/one-prometheus-for-unraid-synology-and-a-vps.md new file mode 100644 index 0000000..f943fc0 --- /dev/null +++ b/src/content/posts/zh/one-prometheus-for-unraid-synology-and-a-vps.md @@ -0,0 +1,125 @@ +--- +title: "用一个 Prometheus 监控 unRaid、群晖 NAS 和 VPS:我做错的四件事" +description: "三台主机、一个 Prometheus、一个 Grafana:覆盖矩阵、把同一个 NAS 卷算成三次的\"总容量\",以及那台完全不导出 CPU 频率的虚拟机。" +pubDate: 2026-07-29 +updatedDate: 2026-09-29 +category: case-studies +tags: [prometheus, grafana, unraid, synology, dsm, cadvisor, monitoring, homelab] +ogImage: /og/one-prometheus-for-unraid-synology-and-a-vps.png +banner: /banners/one-prometheus-for-unraid-synology-and-a-vps.png +draft: false +--- + +## 为什么这件事重要 + +当一台家用服务器在跑别人的网站、邮件和文件时,它就不是玩具了。糟糕的一周和糟糕的**一个月**之间,差别通常只是你多早发现:一块出现重分配扇区的盘、一个只剩 20 GB 的卷、一台被宿主机饿死的虚拟机、一个夜里重启四次的容器。 + +我有三台机器——unRaid(日常主力)、群晖 NAS(存储与服务)和 VPS(对外)——以及三套**各自**都不知道它们在干什么的方式。这篇文章讲它们如何变成一块屏,以及我路上做错的四件事——其中三个你也会踩。 + +## 架构 + +一个 Prometheus 加一个 Grafana,跑在永不关机的那台机器上。每台被监控主机跑两个 exporter: + +| 机器 | 主机指标 | 容器指标 | 抓取路径 | +|---|---|---|---| +| unRaid | node_exporter (`:9100`) | cAdvisor (`:8080`) | 直连 | +| 群晖 DSM | node_exporter (`:9100`) | cAdvisor (`:8082`) | 局域网 | +| ServerHosh VPS | node_exporter (`:9100`) | cAdvisor (`:8081`) | VPN 隧道 + `nginx` stream 中转 | + +**node_exporter 与 cAdvisor 不是替代关系——这是最容易搞错的一点。** node_exporter 看的是**主机**:CPU、内存、网络、磁盘、文件系统、温度;它对 cgroup 一无所知,说不出是哪个容器吃掉了内存。cAdvisor 只看**容器**。想要主机健康和按容器记账,两个都得跑;少一个,那个洞会在几个月后表现为一次无法解释的负载高峰。 + +## 错误一:cAdvisor 里那些\"不是容器的容器\" + +我的容器规则开始对不是容器的东西告警。cAdvisor 会为 **systemd slice** 和整机 cgroup 导出序列——其中一条没有名字的序列报告了约 58 GB 的\"内存使用\",却没有对应的容器。那就是整台主机,被描述成了一个容器。 + +修法是加一个过滤条件,但你得先知道要写它: + +```promql +# container metrics: anything with a name, and only that +container_memory_working_set_bytes{job=~"cadvisor.*", name!=""} +``` + +少了 `name!=""`,一条\"容器内存超过 10 GB\"的规则会对主机自己告警。 + +## 错误二:我的\"总容量\"是虚构的 + +聚合面板好看,但就是错的。原因:**同一个文件系统会在指标里出现多次。** + +- 在 NAS 上,主卷是 `/volume1`,而 `/opt` 是**同一个** btrfs 文件系统,只是换了个挂载点。 +- 在 unRaid 上,同一个 NAS 卷又以 CIFS 挂载成 `/mnt/remotes/_ActiveBackup`。 +- unRaid 的 `/var/lib/docker` 是那个池的子卷,而 `/mnt/ssd` 已经代表过同一个池——同样的字节,第二个身份。 + +所以朴素的 `sum(node_filesystem_size_bytes)` 报出了我并不存在的容量。诚实的写法是明确列出什么算数据存储: + +```promql +node_filesystem_size_bytes{ + mountpoint=~"/volume[0-9]+|/mnt/ssd|/mnt/disk[0-9]+", + fstype!~"fuse.*|tmpfs|rootfs" +} or node_filesystem_size_bytes{job="vps-host", mountpoint="/"} +``` + +还有两条应该写进面板说明的诚实备注,因为没人能解释的数字比没有数字更糟: + +- **校验盘对内核不可见。** unRaid 的奇偶校验盘没有文件系统,所以从不出现在指标里——这个和是可用容量,不是硬盘数量。 +- **镜像会虚增裸和。** 两块镜像 SSD 会各报一次自己的字节,和不是你能存的量。 + +还有那个以后一定会咬我的遗漏:unRaid 的 shfs 联合挂载(`/mnt/user`)报告 `avail=0`。把它算进\"剩余空间低于 X\"的规则里,就会产生一条永远无法解除的告警。 + +## 错误三:那台不导出 CPU 频率的虚拟机 + +为了算全家的\"总 CPU 频率\",我直接用 node_exporter 的 cpufreq collector。unRaid 和 NAS 都按核心报出了 `node_cpu_scaling_frequency_hertz`;VPS **什么都没有**——它是 KVM 客户机,而客户机没有 `/sys/devices/system/cpu/cpu0/cpufreq`。这个指标在那里无法存在。 + +修法是 textfile collector:一个小脚本读客户机**能**看到的东西(`/proc/cpuinfo`),把 Prometheus 格式指标写进 node_exporter 会抓取的目录: + +```sh +# /opt/node-exporter-textfile/cpu-mhz.sh — cron 每 5 分钟 +awk -F: ' + /^processor/ { c = $2; gsub(/[ \t]/, "", c) } + /^cpu MHz/ { f = $2; gsub(/[ \t]/, "", f); printf "node_cpu_mhz_current_hz{core=\"%s\"} %.0f\n", c, f * 1000000 } +' /proc/cpuinfo +``` + +配合: + +```yaml +command: + - '--collector.textfile.directory=/textfile' +volumes: + - /opt/node-exporter-textfile:/textfile:ro +``` + +两个刻意的决定:指标名**故意不同**于 node_exporter 自己的(用 `node_cpu_mhz_current_hz`,不用 `node_cpu_scaling_frequency_hertz`),这样万一以后宿主机暴露真实 cpufreq,不会出现重复序列冲突;仪表盘则显式合并两个来源: + +```promql +sum(node_cpu_scaling_frequency_hertz or node_cpu_mhz_current_hz) +``` + +## 错误四:exporter 把容器名当成了主机名 + +仪表盘的主机下拉框里出现了 `9f9afcccc962`。那是个容器 ID:node_exporter 的 `uname` collector 读到的是**进程自己的** UTS namespace,而容器的 hostname 默认就是它自己的 ID。指标没错,标签毫无意义。compose 里一行解决: + +```yaml +services: + node_exporter: + hostname: 2.hoelee.com # otherwise `nodename` = the container ID +``` + +## 我下次会怎么做 + +- **刻意把总览屏留到最后做。** \"总容量是多少\"这个问题会逼出全部去重工作;先做完三台主机面板再发现,就得在每个地方重做这个数字。 +- **一开始就给容器钉 hostname。** 一行成本,省掉一个让人困惑的下拉框。 +- **假设每个虚拟化或家电式主机都会藏掉一类指标。** NAS 可能限制你的容器监控(DSM 的 Docker API 版本把我的 cAdvisor 钉在 v0.53.0——更新的版本要更新的 Docker API);虚拟机藏掉 cpufreq;路由器藏掉自己的 CPU。找缺口的方法是**数你期待的东西**,而不是相信\"采集器成功了\"。 + +## 结果 + +三台主机、**7 个抓取目标、一个 Grafana、16 条告警规则**,以及一块屏:**47 个 CPU 核心 · 当前 158 GHz(标称上限 205 GHz)· 142 GB 内存 · 64 TB 存储 · 155 个运行中容器**,并且每台机器克隆同一套主机/容器面板布局,三台读起来完全一致。告警进 Telegram 和邮件,Prometheus 保留 30 天。 + +最有用的产出不是仪表盘,而是重写存储规则那天触发的一条告警:一个只剩 21 GB 的 NAS 卷——它一直藏在\"99% 满\"的百分比阈值背后,而那个巨物早已变成背景噪声。 + +## 需要为你的业务做这个吗? + +如果你有 NAS、VPS 和几台服务器,却没有一个地方能同时看它们,我可以帮你搭这样的自托管监控流水线——Prometheus + Grafana 覆盖你的机器,主机**与**容器指标,真正有意义的磁盘健康与容量规则,通知推到 Telegram 或邮件。 + +**WhatsApp:[+60 12-797 2969](https://wa.me/60127972969)** · **邮箱:[me@hoelee.com](mailto:me@hoelee.com?subject=Self-hosted%20monitoring%20setup)** · **[hoelee.com](https://hoelee.com)** + +网站设计与开发是我的主业;服务器加固与自托管基础设施是它的另一半。 diff --git a/src/content/posts/zh/smartctl-exit-code-32-skips-the-disks-that-matter.md b/src/content/posts/zh/smartctl-exit-code-32-skips-the-disks-that-matter.md new file mode 100644 index 0000000..40cfab6 --- /dev/null +++ b/src/content/posts/zh/smartctl-exit-code-32-skips-the-disks-that-matter.md @@ -0,0 +1,96 @@ +--- +title: "smartctl 退出码 32:专门跳过你最该看的硬盘的那个\"错误\"" +description: "把 smartctl 的任何非零退出码当成读取失败的监控脚本,恰好会跳过属性已经逼近阈值的那些盘。退出码 32 不是错误,而是一段历史。" +pubDate: 2026-04-14 +updatedDate: 2026-09-29 +category: notes +tags: [smart, smartctl, unraid, monitoring, bash, disks, homelab] +ogImage: /og/smartctl-exit-code-32-skips-the-disks-that-matter.png +banner: /banners/smartctl-exit-code-32-skips-the-disks-that-matter.png +draft: false +--- + +我的硬盘健康采集脚本连续几周都是\"全绿\"。每块盘都有温度、通电小时数和 SMART 结论,指标看起来完整——直到我数了一下只有 **4 块 SSD 里的 2 块**。 + +不是故障,是**消失了**。脚本认定这两块盘不存在。 + +## 问题出在把退出码当布尔值 + +采集脚本遍历块设备、对每块盘跑 `smartctl`、再写 Prometheus textfile 指标。错误处理看起来挺防御: + +```bash +for dev in /dev/sd?; do + if ! smartctl -A -H "$dev" > /tmp/smart.out 2>&1; then + continue # "盘不支持 SMART / 读不到" + fi + # 解析并输出指标 +done +``` + +`if ! cmd` 是布尔判断,但 `smartctl` 的退出状态**不是布尔值,是一个位域**。把位域当真假用,就是这次翻车的原因。 + +## 退出码 32 的真实含义 + +`man smartctl` 里这些位是独立累加的: + +| 位 | 值 | 含义 | +|---|---|---| +| 0 | 1 | 命令行解析失败 | +| 1 | 2 | 设备打开失败 / 无 IDENTIFY DEVICE 结构 | +| 2 | 4 | SMART 或 ATA 命令失败、校验和错误 | +| 3 | **8** | SMART 状态为 **DISK FAILING** | +| 4 | **16** | 有预失效属性 **<= 阈值** | +| 5 | **32** | SMART 状态 **OK**,但某些属性**曾经**低于阈值 | +| 6 | 64 | 设备错误日志中有记录 | +| 7 | 128 | 自检日志中有失败记录 | + +所以 `32` 的意思和\"读不到盘\"正好相反:**现在没问题,但它曾经踩过阈值。** 这正是应该持续盯着的那类盘——而我的脚本把它们丢掉了。 + +消失的两块盘,恰好是原始属性长期处在边缘值的那两块;退出码为 `0` 的\"健康盘\"被正常采集。**这个过滤器实际上筛选出了\"没有任何历史可报告\"的盘。** + +## 修法:把\"信息位\"掩掉 + +对监控来说,位 32 和 64 是信息;位 8 和 16 才是该告警的。掩掉信息位,只对剩下的做判断: + +```bash +smartctl -A -H -d sat "$dev" > /tmp/smart.out 2>&1 +rc=$? + +# fatal bits: 1 (parse), 2 (open), 4 (command/checksum), 8 (FAILING) +# informational bits: 32 (was below threshold in the past), 64 (error log has records) +fatal=$(( rc & ~(32 | 64) )) +if [ "$fatal" -ne 0 ]; then + echo "device $dev unreadable or failing (rc=$rc)" >&2 + continue +fi + +# emit the verdict AND the exit code, so the code itself is a metric +echo "disk_smart_exit_code{device=\"$dev\"} $rc" +echo "disk_smart_health{device=\"$dev\"} $(( rc & 8 ? 0 : 1 ))" +``` + +三点改变: + +1. **退出码变成数据,而不是控制流。** 它作为指标被导出,磁盘从 `0` → `32` → `64` 的漂移会变成趋势,而不是一次静默跳过。 +2. **信息位不再致命。** 有历史的盘被采集,这才是监控它们的目的。 +3. **NAS 上 `-d sat` 很关键。** 在 Synology DSM(以及某些 USB 桥接)上,设备类型不对时 `smartctl` 拿不到有用输出——症状同样是\"盘消失\",原因却不同。 + +## 我下次会怎么做 + +- **面对有\"退出码位域\"文档的工具,永远不要写 `if ! cmd`。** 先把手册里的退出码那节读完。 +- **除了\"采集器有没有跑\",还要数\"采到了几块盘\"。** 我的告警是\"采集器过期\",真正的 bug 是\"采集器正常跑了,只报了 50% 的盘\"。一个 `disk_count` 指标就能立刻暴露。 +- **把\"这个设备没有数据\"本身当成可告警状态。** 缺失的序列是不可见的,所以它存活了好几周。 + +## 结果 + +采集器从 2 块可用盘变成 4 块,每块都报温度、通电小时数、SMART 状态和原始退出码——共 122 个 textfile 指标,包括我真正想要的 btrfs 错误计数。重新出现的那两块,正是长期处在边缘属性值的盘。 + +一个会跳过自己最坏信号的监控流水线,比没有监控更糟——因为它一直告诉你一切正常。 + +## 需要为你的业务做这个吗? + +如果你在跑 NAS 或服务器机架,想让硬盘健康真正\"会通知你\"——SMART 属性、温度、btrfs/RAID 错误计数,并推到 Telegram 或邮件——这类自托管监控流水线我可以帮你搭。 + +**WhatsApp:[+60 12-797 2969](https://wa.me/60127972969)** · **邮箱:[me@hoelee.com](mailto:me@hoelee.com?subject=Disk%20health%20monitoring)** · **[hoelee.com](https://hoelee.com)** + +网站设计与开发是我的主业;服务器加固与自托管基础设施是它的另一半。 diff --git a/src/content/posts/zh/why-your-grafana-dashboard-shows-no-data.md b/src/content/posts/zh/why-your-grafana-dashboard-shows-no-data.md new file mode 100644 index 0000000..1d4f16b --- /dev/null +++ b/src/content/posts/zh/why-your-grafana-dashboard-shows-no-data.md @@ -0,0 +1,112 @@ +--- +title: "Prometheus 明明有数据,Grafana 面板却全是 No data 的原因" +description: "所有 target 都是 up、数据就在 Prometheus 里、面板表达式也没错——41 个面板却全部 No data。原因是 Grafana 变量过滤了自己,而我的验证方式恰好掩盖了它。" +pubDate: 2026-05-17 +updatedDate: 2026-09-29 +category: devops +tags: [grafana, prometheus, dashboards, monitoring, promql, observability] +ogImage: /og/why-your-grafana-dashboard-shows-no-data.png +banner: /banners/why-your-grafana-dashboard-shows-no-data.png +draft: false +--- + +打开每天在看的面板,结果是空的。不是\"某一格坏了\",而是**从上到下每一格**都写着 *No data*。去查 Prometheus,一切健康。这篇讲我踩到的那个具体且不直观的原因,以及为什么我自己的验证会同时告诉我\"面板没问题\"。 + +## 最省时间的排查顺序 + +常见原因(数据源选错、时间范围不对、`rate()` 用在会重置的计数器上、target 没被抓取)都不适用。真正有效的检查顺序是: + +```bash +# 1. target 真被抓取了吗? +curl -s http://prometheus:9090/api/v1/targets | jq '.data.activeTargets[] | {job:.labels.job, health}' + +# 2. 指标现在存在吗?(完全绕过 Grafana) +curl -s -G http://prometheus:9090/api/v1/query \ + --data-urlencode 'query=node_uname_info{job="unraid-host"}' | jq '.data.result[0].metric' + +# 3. 面板表达式本身对吗? +curl -s -G http://prometheus:9090/api/v1/query \ + --data-urlencode 'query=count(node_cpu_seconds_total{mode="idle",job="unraid-host"})' +``` + +三项都过了:7 个 target `up`、2.5 万条序列在采集、手动跑面板表达式也有数据。所以问题不在指标、不在抓取、也不在 PromQL——而在**它们之间的变量层**。 + +## 真正的原因:变量过滤了自己 + +面板上有一个主机选择器,两个变量串联:`$nodename`(哪台机器)和 `$node`(它的 `instance` 标签)。第一个的定义是: + +```promql +# BROKEN — 变量用自身的值来过滤自己 +label_values(node_uname_info{job="unraid-host", nodename=~"$nodename"}, nodename) +``` + +首次加载时 `$nodename` 还没有值,选择器就变成 `nodename=~""`——**什么都匹配不到**。查询返回零行的变量没有任何选项,于是它保持空;而 `$node` 又是基于 `$nodename` 定义的: + +```promql +label_values(node_uname_info{job="unraid-host", nodename="$nodename"}, instance) +``` + +于是 `$node` 也是空,所有以 `$node` 为过滤条件的面板自然匹配不到序列。Grafana 说 *No data* 并没有骗人:面板查询真的没有数据,因为给它限定范围的变量解析成了空。 + +修法是去掉自引用,并给每个变量存一个明确的默认值: + +```promql +# FIXED — 不自引用,每个变量只跳一层 +# $nodename +label_values(node_uname_info{job="unraid-host"}, nodename) +# $node +label_values(node_uname_info{job="unraid-host", nodename="$nodename"}, instance) +``` + +## 为什么我的验证没抓到 + +这才是值得抄的部分。我\"验证\"面板的方式是:**手动**把每个面板表达式里的变量替换成我知道正确的值(`$nodename=unRaid`、`$node=host.docker.internal:9100`),然后断言查询有数据点。13 项检查全部通过——面板依然是空的,因为 bug 在被替换的那个值**上游**:Grafana 自己对变量的解析是空的。 + +**一个替换掉\"可疑值\"的验证,永远发现不了那个值解析失败。** 要抓它,得读 Grafana 真实存下来的值: + +```bash +curl -s -u admin:"$PW" http://grafana:4010/api/dashboards/uid/rYdddlPWk \ + | jq '.dashboard.templating.list[] | {name, current: .current.value}' +``` + +更直接的检查对象是变量查询本身——跑 Grafana 会跑的那条: + +```bash +# $nodename 的 label_values() 到底返回什么? +curl -s -G http://prometheus:9090/api/v1/query \ + --data-urlencode 'query=count by (nodename) (node_uname_info{job="unraid-host"})' | jq '.data.result' +``` + +那里是空,面板就不可能渲染出来,无论数据多健康。 + +## 第二个坑:`$__all` 不是 `.*` + +做自动化面板检查时,值为\"All\"的变量,其 current 是哨兵字符串 `$__all`——不是正则。把 `name=~"$__all"` 原样展开会匹配不到任何东西,于是完全健康的容器面板\"全军覆没\"。展开前要把 `$__all` 解析成该变量的 `allValue`(通常是 `.*`),否则你追的是一个只存在于自己检查脚本里的 bug。 + +## 不是每个空面板都是 bug + +修完之后 25 个面板有 21 个出数。剩下 4 个是**设计上就不可能**有数据的,最好在面板说明里写清楚: + +- **PSI 面板**需要 `/proc/pressure`,而这台 NAS 的内核不暴露它——指标在这台机器上无法存在;同样的面板在会导出 `node_pressure_*` 的主机上正常显示。 +- **根文件系统面板**的表达式带 `fstype!="rootfs"`,而某台主机的 `/` 是**内存盘**(unRaid 就是),所以永远匹配不到。 + +区分\"因为 bug 空\"和\"因为这台机器不可能有这个指标而空\",是真修好和瞎忙一下午的分界线。 + +## 我下次会怎么做 + +- **永远不要定义过滤自己的模板变量。** 一个变量只跳一层。 +- **给面板依赖的变量存明确的 `current` 值**,别让全新加载依赖下拉框被选中。 +- **用 Grafana 存下来的值做验证**,而不是我以为的值。 +- **故意留空的缺口写进面板说明**,否则未来的你会花一个下午去\"修\"一个本来就没打算工作的面板。 + +## 结果 + +三台主机现在跑同一套布局——每台一个 41 面板的主机面板 + 一个 10 面板的容器面板,分别有 21/23/25 个面板出数,两类结构性空缺写在面板里,而不是留给下一个打开它的人去猜。 + +## 需要为你的业务做这个吗? + +如果你有一堆没人信任的 Grafana 面板,或者服务器和 NAS 完全没有监控,我可以帮你搭自托管的 Prometheus + Grafana(主机与容器指标、合理的告警规则、通知推到 Telegram 和邮件),也会把你那些悄悄\"不出数\"的面板修好。 + +**WhatsApp:[+60 12-797 2969](https://wa.me/60127972969)** · **邮箱:[me@hoelee.com](mailto:me@hoelee.com?subject=Grafana%20monitoring)** · **[hoelee.com](https://hoelee.com)** + +网站设计与开发是我的主业;服务器加固与自托管基础设施是它的另一半。 diff --git a/src/content/posts/zh/your-disk-full-alert-is-lying.md b/src/content/posts/zh/your-disk-full-alert-is-lying.md new file mode 100644 index 0000000..b3334d0 --- /dev/null +++ b/src/content/posts/zh/your-disk-full-alert-is-lying.md @@ -0,0 +1,87 @@ +--- +title: "你的\"磁盘将满\"告警在说谎:别再按百分比告警" +description: "\"剩余 3%\"听着像紧急事故,直到你发现那是 500 GB。在多 TB 的卷上,百分比阈值会在没事的时候乱叫,也会在真出事的时候沉默。" +pubDate: 2026-06-24 +updatedDate: 2026-09-29 +category: devops +tags: [monitoring, alerting, prometheus, grafana, storage, capacity, observability] +ogImage: /og/your-disk-full-alert-is-lying.png +banner: /banners/your-disk-full-alert-is-lying.png +draft: false +--- + +我搭好监控、写了一套看着合理的规则,然后用一个下午删掉了其中大半阈值。它们全是百分比,而且都以同一种方式错了。 + +## 三个教我道理的告警 + +**1.「文件系统使用率超过 90%」**——它在一个还剩约 500 GB 的 17 TB 卷上触发。没事发生,也不会马上有事发生。这种体量的卷,一次大备份就能让百分比摆动几个点:比例是噪声,而\"还有 500 GB 可用\"才是世界的真实状态。 + +**2.「内存使用率超过 90%」**——在一台**可用**内存 4.8 GB 的主机上反复触发。Linux 会用空闲内存做 page cache,需要时立刻归还;`MemTotal - MemFree` 描述的是文件缓存,不是你的风险。真正预示 OOM 的是 `MemAvailable`。 + +**3.「CPU steal 超过 25%」**——一台把邮件栈跑得完全正常的 VPS,steal 长期在 41%。steal 说明宿主机忙,它是**容量**信号而不是**故障**信号;7 vCPU 的机器 41% steal 就是这台机器的成本。这条规则产生的告警永远为真、也永远无用。 + +共同点:**我在对比例告警,而比例不知道东西有多大。** + +## 对\"后果\"告警,而不是对比例 + +我现在对每个阈值都问一句:*如果这个状态持续下去,实际会发生什么?* 答案几乎总能用绝对单位表达。 + +| 不要这样 | 改成 | 原因 | +|---|---|---| +| 文件系统 > 90% | 剩余空间 < N GB | \"我还能不能写进去\"才是问题 | +| 内存使用 > 90% | `MemAvailable` < 512 MB | 逼近 OOM,而不是\"缓存很暖和\" | +| CPU steal > 25% | steal > 50% | 容量成本 vs 真的被饿死 | +| 负载 > N | `load1 / 核数 > 2` | 不谈核数的负载没有意义 | + +PromQL 的前后对比: + +```promql +# BEFORE — 17 TB 卷的百分比;还剩半 TB 就开叫 +(1 - node_filesystem_avail_bytes / node_filesystem_size_bytes) * 100 > 90 + +# AFTER — 真正装数据的卷上,还剩多少容量 +min by (instance, mountpoint) ( + node_filesystem_avail_bytes{mountpoint=~"/volume[0-9]+|/mnt/ssd|/mnt/disk[0-9]+"} +) < 25e9 +``` + +```promql +# BEFORE — 把 page cache 当成压力 +(1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100 > 90 + +# AFTER — 真正的 OOM 前兆 +node_memory_MemAvailable_bytes < 512 * 1024 * 1024 +``` + +两个让规则能活下来的细节: + +**把比较放进阈值条件里,不要塞进表达式。** 查询就写 `min by (mountpoint) (node_filesystem_avail_bytes{...})`,让告警规则去和 `25000000000` 比。若你在 PromQL 里写死 `< 25e9`、规则里又留着继承来的 `> 90`,规则照样会评估——但界面显示的阈值是错的,下一个读它的人得从两个互相矛盾的条件下反推你的意图。 + +**排除那些\"只是另一个挂载点的视图\"。** 一条\"剩余空间\"规则有多诚实,取决于它的选择器。在我自己的环境里,同一个 NAS 卷出现三次(`/volume1`、`/opt`、以及在另一台主机上以 CIFS 再次挂载),而 unRaid 的 shfs 联合挂载报告 `avail=0`——不加排除的规则不是永远误报,就是把容量算成两倍。一个写明白的选择器胜过聪明的通用写法。 + +## 噪声不是无害的 + +坏阈值的代价不在告警本身,而在**训练**:每条\"没事乱叫\"的告警都在教你先扫一眼、然后忽略;等你真正该看的那条到来时,它出现在一个你早已不再读的频道里。我的规则总数**减少了**,覆盖率却上升了——更少,但每条都对应一个明确的后果。 + +## 阈值是私人的,形状是通用的 + +上面的数字是我的:NAS 卷剩不到 25 GB 我会在意、可用内存低于 512 MB 我才管、VPS steal 超过 50% 才算被饿。你的数字随硬件和容忍度而变。 + +能通用的是形状: + +- 用绝对单位而非比例; +- 把后果写进告警的摘要里; +- 选择器明确声明它在替哪些挂载点说话; +- 规则数量少到你还能读完每一条。 + +## 结果 + +同一套栈,从\"三条永久无用报警\"变成 16 条在机房健康时保持安静的规则——其中一条立刻暴露了一个只剩 21 GB 的 NAS 卷,而那正是百分比规则一直藏在\"99% 满的巨物\"背后、被你学会忽略的东西。 + +## 需要为你的业务做这个吗? + +如果你的监控在发你已经学会忽略的告警,或者你根本没有监控、宁愿从一条规则而不是从一次失败的备份里得知磁盘满了,我可以帮你搭自托管的监控与告警(Prometheus + Grafana,通知推 Telegram 和邮件),阈值按\"什么会真的坏\"来定。 + +**WhatsApp:[+60 12-797 2969](https://wa.me/60127972969)** · **邮箱:[me@hoelee.com](mailto:me@hoelee.com?subject=Monitoring%20and%20alerting)** · **[hoelee.com](https://hoelee.com)** + +网站设计与开发是我的主业;服务器加固与自托管基础设施是它的另一半。