post(devops): it reported success, nothing had changed (EN+ZH) + og/banner

Three silent failures on a network-mounted share: a 403 MB upload that wrote 12.5 MB and returned OK
(the timeout toast arrived minutes later, asynchronously), an os.replace() that never landed while
getsize()/mtime served cached metadata (the destination was still the original H.264 file), and a
byte-count source check that agreed with a length-preserving bad copy. Adds read-back hashing at all
three boundaries (source copy, upload, install) and records the metric traps that hid them: a clean
decode test on wrong pictures, SSIM on near-flat frames, input-seek decoding of open-GOP sources,
stream-copy slices dropping frames, and fixed-name scratch files colliding across concurrent runs.
Audit after the fix: reported=30 landed=29 not-landed=0.
This commit is contained in:
2026-10-08 05:04:54 +08:00
parent e6476e1d15
commit 0b523f40d6
6 changed files with 296 additions and 0 deletions
Binary file not shown.

After

Width:  |  Height:  |  Size: 82 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 39 KiB

+18
View File
@@ -1167,6 +1167,24 @@ BANNERS['every-player-stuttered-the-file-was-fine'] = {
],
};
BANNERS['it-reported-success-nothing-had-changed'] = {
titlebar: "root@nas - reported success",
lines: [
{ t: 'prompt', text: "$" }, { t: 'cmd', text: "convert-batch --tier x265" },
{ t: 'dim', text: "stage: 403,448,100 bytes written" }, { t: 'err', text: "remote: 12,582,912 bytes - reported OK" },
{ t: 'prompt', text: "$" }, { t: 'cmd', text: "read-back hash -> mismatch, re-upload" },
{ t: 'ok', text: "hash match - replacing" }, { t: 'hl', text: "verify at every boundary" },
{ t: 'dim', text: "reported=30 landed=29 not-landed=0" }, { t: 'dim', text: "one writer, zero silent failures" },
],
flow: [
{ n: '1', label: "stage" },
{ n: '2', label: "upload" },
{ n: '3', label: "re-read" },
{ n: '4', label: "hash" },
{ n: '5', label: "replace" },
],
};
// ---------- read frontmatter ----------
const postPath = join(ROOT, 'src', 'content', 'posts', `${slug}.md`);
let category = 'devops';
+6
View File
@@ -355,6 +355,12 @@ TERMINALS['every-player-stuttered-the-file-was-fine'] = `
<div class="line"><span class="prompt">$</span><span class="cmd">-c:v av1_cuvid -> -hwaccel cuda -c:v av1</span></div>
<div class="line"><span class="prompt">&nbsp;</span><span class="cmd">0 wrong frames SSIM 1.000000</span><span class="fix">-> fixed</span></div>`;
TERMINALS['it-reported-success-nothing-had-changed'] = `
<div class="line"><span class="prompt">$</span><span class="cmd">upload show.mp4 -> "ok"</span></div>
<div class="line"><span class="prompt">&nbsp;</span><span class="err">local 403,448,100 B / remote 12,582,912 B</span></div>
<div class="line"><span class="prompt">$</span><span class="cmd">read-back hash -> mismatch</span></div>
<div class="line"><span class="prompt">&nbsp;</span><span class="cmd">re-upload -> hash match</span><span class="fix">-> verified</span></div>`;
// ---------- read frontmatter ----------
const postPath = join(ROOT, 'src', 'content', 'posts', `${slug}.md`);
if (!existsSync(postPath)) {
@@ -0,0 +1,169 @@
---
title: "It Reported Success. Nothing Had Changed."
description: "Three ways a network-mounted drive lied to my conversion pipeline: a write that silently stopped at 12.5 MB of 403 MB, a rename that never landed while the metadata said otherwise, and a size check that agreed with a corrupted copy."
pubDate: 2026-10-08
category: devops
tags: [storage, smb, verification, ffmpeg, python]
ogImage: /og/it-reported-success-nothing-had-changed.png
banner: /banners/it-reported-success-nothing-had-changed.png
draft: false
---
## The batch that looked fine
I run a conversion pipeline over a 582-file video library that lives on a
NAS and is mounted as a Windows drive letter through a third-party SMB
client (the kind that presents the share as a normal `Z:\` drive). Small
files, big media files, one writer.
The pipeline does the obvious things: stage the source locally, encode,
verify the output, upload it next to the original, replace the original. It
logged success on every step I had thought to check.
Then I started auditing the *results* instead of the log. Over a set of 30
files the pipeline had reported as done, **one had not changed at all** —
the file on the share was still the original. No error, no warning, no
failed step. The log said `ok`.
That is the failure mode this post is about: not a crash, not a corrupted
file, but a system that reports an outcome it did not achieve. If your
pipeline's success is defined by its own log, you have no idea what is
actually on disk.
## Lie 1: a write that stops early and returns success
The first one is the scariest because everything looks right at the moment
of the call.
```text
FAIL read-back mismatch
local = 3b7cd25d1de1 403,448,100 bytes
remote = 6a7785351031 12,582,912 bytes
```
The upload of a 403 MB output file wrote **12.5 MB** and returned. No
exception, no non-zero exit, no short-write return code. The client's own
error — a timeout toast — appeared **asynchronously, minutes later**, long
after my code had moved on and printed its success line.
This is why a "did the write succeed?" check that trusts the API is worth
nothing on this class of storage. The write call answered a different
question ("did I hand the data to the cache?") than the one you asked ("is
the file on the other side correct?").
The fix is boring and total: hash what you wrote locally, read the file back
from the share, hash that, and compare. On a mismatch, re-upload. Twice, if
needed, with a backoff.
```python
# after uploading dst from src
if hash_file(src) != hash_file(_reread(dst)):
retry_upload()
```
Note the deliberate cost: this reads the file over the network a second
time. For a 400 MB file that is a few seconds — against a 25-minute encode
it is rounding error. The alternative is a library where some fraction of
files are silently not what you think.
## Lie 2: a rename that does not land, with metadata that says it did
The second one hides behind a filesystem primitive you trust blindly:
`os.replace()` — atomic, same-filesystem, cannot half-happen. Except
"filesystem" here means a network client with its own metadata cache.
The sequence was: upload `show.mp4.new.mp4`, verify the hash, `os.replace()`
it over `show.mp4`, then probe the result to confirm. The probe said HEVC,
the size looked right, the run logged `ok 1085 MB -> 478 MB`.
The file on the share was **the untouched original**: H.264, 921,950,850
bytes. The replace had not been applied; the client served cached attributes
for the path, so `getsize()` and `mtime` happily described a file that no
longer existed.
The only check that survives a lying metadata cache is reading the bytes
back:
```python
os.replace(tmp, dst) # claim
if hash_file(dst) != hash_file(local): # evidence
raise VerificationError(dst)
```
After I added that step, the same audit that found 1-in-30 failures found
**zero** — and, more importantly, any future silent failure turns into a
loud one, because "the bytes at the destination do not match the bytes I
produced" is not a thing a cache can fake.
## Lie 3: a size check that agrees with a bad copy
The third one is the subtlest, because it fools *verification itself*, not
just the operation.
Staging a source file from the share into local scratch is a copy. The
natural check is "did the right number of bytes arrive?" — so the pipeline
compared byte counts. They matched.
Then, in an unrelated investigation, the same source file failed to decode
("Error splitting the input into NAL units"), and I went looking for a
corrupt file that was not there. The file was fine: 921,950,850 bytes,
healthy header, sixty seconds decoded with zero errors. The *copy* had been
truncated in a way that preserved the length, or the read had silently
returned short and been padded.
Length is not a checksum. A copy is verified when the hash of what you read
matches the hash of what is on the share — which means reading the share
twice, or hashing after the read and comparing against a hash taken on a
different occasion.
## The metric traps that made it worse
While chasing those three, my *quality* checks were lying too, in ways worth
naming because they are not specific to network storage:
- **A decode test passes on wrong pictures.** `ffmpeg -v error -i out.mp4 -f
null -` printed nothing on a file with 106 frames decoded from the wrong
timestamps. "No errors" means "no errors", not "correct".
- **SSIM is meaningless on flat frames.** Two nearly-black frames can score
0.25 while being visually identical. Any threshold that treats a single low
frame as proof of a defect will reject good encodes — and any that ignores
frames wholesale will miss real ones.
- **Seeking into an open-GOP source decodes the wrong pictures.** Comparing a
file against *itself*, with one side decoding from a mid-file seek and the
other sequentially, measured min SSIM 0.25. A false "systematic defect"
produced entirely by my own comparison harness.
- **A stream copy is not a lossless slice.** Cutting 60 seconds out of an
H.264 file with B-frames using `-ss ... -c copy` produced 1270 frames where
1500 were expected. Fine for scrubbing; fatal as a test fixture.
- **Two concurrent runs sharing a scratch directory will overwrite each
other's intermediates.** Mine had fixed filenames for the decoded raw
frames, so two gates running at once compared two different videos and
reported 100% of frames below the threshold. Every temporary path now
carries the process id.
## The discipline
None of this needs cleverness. It needs a rule: **a step is done when you
have independently observed the outcome, not when the tool reported it.**
In practice, for a pipeline that writes to a share:
1. **Read the source twice.** Stage the file, hash it, hash a fresh read of
the share, compare. Length is not evidence. 2. **Hash the upload before you
replace anything.** 3. **Hash the destination after the replace.** Cached
metadata cannot fake bytes. 4. **Verify the content, not just the
container** — compare pictures, not frame rates. 5. **Make every failure
loud.** A gate that fails safe costs one wasted encode; a gate that passes a
bad file costs a corrupted library and your trust in it. 6. **Do not
parallelise the writer.** Three jobs sharing one share and one scratch
directory produced short writes, timeouts and cross-contaminated
comparisons. When I dropped to one writer, the error rate went to zero.
The audit line I care about is not "0 errors". It is:
```text
reported=30 landed=29 not-landed=0 wrong-codec=1 gone=0
```
Every number there was produced by reading the destination, not by trusting
the log.
@@ -0,0 +1,103 @@
---
title: "它报告了成功,可实际上什么都没变"
description: "网络挂载盘对我的转换流水线撒的三种谎:一次 403 MB 的写入悄悄停在 12.5 MB、一次从未真正落地的重命名(元数据却言之凿凿)、以及一个和坏拷贝达成一致的字节数检查。"
pubDate: 2026-10-08
category: devops
tags: [storage, smb, verification, ffmpeg, python]
ogImage: /og/it-reported-success-nothing-had-changed.png
banner: /banners/it-reported-success-nothing-had-changed.png
draft: false
---
## 一批"看起来没问题"的任务
我跑着一条转换流水线,处理一个 582 个文件的视频库。库在 NAS 上,通过第三方 SMB 客户端挂载成 Windows 盘符(就是那种把共享目录呈现成普通 `Z:\` 盘的东西)。文件有大有小,只有一个写入者。
流水线做的都是些显而易见的事:把源文件暂存到本地、编码、验证输出、上传到原文件旁边、替换原文件。我想到要检查的每一步,它都记了成功。
然后我开始审计**结果**,而不是审计日志。在 30 个被报告为"已完成"的文件里,**有一个根本没有变**——共享盘上还是原文件。没有报错、没有警告、没有失败步骤。日志写着 `ok`。
这就是本文要讲的失败模式:不是崩溃,不是文件损坏,而是一个**报告了它并未达成之结果的系统**。如果你的流水线用"自己的日志"来定义成功,那你对自己的磁盘上到底有什么,一无所知。
## 谎言一:写入提前停止,却返回成功
第一个最吓人,因为在调用的那一刻一切看起来都对。
```text
FAIL read-back mismatch
local = 3b7cd25d1de1 403,448,100 bytes
remote = 6a7785351031 12,582,912 bytes
```
一次 403 MB 输出文件的上传写入了 **12.5 MB** 就返回了。没有异常、没有非零退出码、没有短写返回码。客户端自己的错误——一个超时提示——是**几分钟后异步弹出**的,那时我的代码早就往下走了,成功日志都打完了。
这就是为什么在这类存储上,一个"信任 API 返回"的写入成功检查一文不值。写入调用回答的是另一个问题("我把数据交给缓存了吗?"),而不是你问的那个("对面那个文件对吗?")。
修法很无聊但很彻底:把你本地写出的东西做哈希,再把共享盘上的文件读回来做哈希,然后比对。不一致就重传——必要时两次,带退避。
```python
# 把 src 上传为 dst 之后
if hash_file(src) != hash_file(_reread(dst)):
retry_upload()
```
注意这个代价是故意的:它要再从网络上把文件读一遍。对一个 400 MB 的文件来说是几秒——相对一次 25 分钟的编码属于舍入误差。而换来的是:不必接受一个"有某个比例的文件悄悄不是你以为的样子"的库。
## 谎言二:重命名没有落地,元数据却说落地了
第二个藏在一个你无条件信任的文件系统原语后面:`os.replace()`——原子的、同一文件系统内的、不可能做一半的。只不过这里的"文件系统"其实是一个自带元数据缓存的网络客户端。
当时的顺序是:上传 `show.mp4.new.mp4`、校验哈希、`os.replace()` 覆盖 `show.mp4`、然后探测结果确认。探测说 HEVC,大小看着对,日志写了 `ok 1085 MB -> 478 MB`。
共享盘上那个文件**是原封未动的原文件**:H.264、921,950,850 字节。替换根本没被应用;客户端为该路径返回了缓存的属性,于是 `getsize()` 和 `mtime` 兴高采烈地描述着一个已经不存在的文件。
唯一能在撒谎的元数据缓存下活下来的检查,是把字节读回来:
```python
os.replace(tmp, dst) # 声称
if hash_file(dst) != hash_file(local): # 证据
raise VerificationError(dst)
```
加上这一步之后,同一次审计在原本 30 个里找出 1 个失败的规模上,找出了 **0 个**——更重要的是,今后任何静默失败都会变成响亮的失败,因为"目标处的字节和我产出的字节不一致"不是缓存能造假的东西。
## 谎言三:一个和坏拷贝达成一致的字节数检查
第三个最隐蔽,因为它骗过的是**验证本身**,而不只是操作。
把源文件从共享盘暂存到本地临时目录是一次拷贝。自然的检查是"字节数对不对?"——所以流水线比了字节数。它们一致。
然后在一次无关的排查里,同一个源文件解码失败了(`Error splitting the input into NAL units`),我就去找一个并不存在的坏文件。那个文件是好的:921,950,850 字节、头部健康、前 60 秒零错误解码。是那份**拷贝**被以某种保持长度不变的方式截断了,或者那次读取悄悄读短了、又被补齐了。
长度不是校验和。一次拷贝只有在"你读到的内容的哈希"与"共享盘上那份的哈希"一致时才算被验证——这要么意味着读两遍,要么意味着读完后算哈希、再和另一时刻取的哈希比对。
## 让事情更糟的那些度量陷阱
追这三个问题的过程中,我的**质量检查**也在撒谎,值得点名,因为它们并不只属于网络存储:
- **解码测试会在错误的画面上通过。** `ffmpeg -v error -i out.mp4 -f null -` 在一个带着 106 帧错误时间戳画面的文件上什么都没打印。"没有报错"只意味着"没有报错",不意味着"正确"。
- **SSIM 在平坦帧上没有意义。** 两帧几乎全黑也能算出 0.25,而它们在视觉上完全相同。任何"单帧过低即判缺陷"的阈值都会拒绝好编码;任何"整体忽略这些帧"的做法又会漏掉真缺陷。
- **对 open GOP 的源做输入 seek,会解出错误的画面。** 把一个文件和**它自己**比对——一侧从中间 seek 进去解码、另一侧顺序解码——测出最低 SSIM 0.25。一个完全由我自己的比对工具制造的"系统性缺陷"假象。
- **流拷贝不是无损切片。** 用 `-ss ... -c copy` 从带 B 帧的 H.264 文件里切出 60 秒,本该 1500 帧却只得到 1270 帧。拿来做拖拽预览可以;拿来做测试夹具是致命的。
- **两个并发任务共用临时目录,会互相覆盖中间产物。** 我的解码裸帧文件用了固定文件名,于是两个同时运行的闸门比的是两个不同的视频,报出"100% 的帧低于阈值"。现在每一条临时路径都带上进程 id。
## 纪律
这些都不需要什么聪明办法。它需要一条规则:**一个步骤算完成,是在你独立观察到了结果之后,而不是在工具报告了结果之时。**
落到一条会往共享盘写东西的流水线上:
1. **源文件读两遍。** 暂存、哈希、再抓一次共享盘的读数做哈希、比对。长度不是证据。
2. **替换任何东西之前,先给上传做哈希。**
3. **替换之后,给目标做哈希。** 缓存的元数据造不了假字节。
4. **验证内容,而不只是验证容器**——比画面,不比帧率。
5. **让每一次失败都响亮。** 一个失败得安全的闸门,代价是白烧一次编码;一个放行了坏文件的闸门,代价是整个库和你对它的信任。
6. **不要把写者并行化。** 三个任务共用一块共享盘和一个临时目录,制造出了短写、超时和互相污染的比较。当我降到单个写者,错误率归零。
我最在意的审计行不是"0 errors",而是:
```text
reported=30 landed=29 not-landed=0 wrong-codec=1 gone=0
```
那里的每一个数字都是**读目标读出来的**,不是信任日志得来的。