Add case study: AI furniture compositing with FLUX.1 Kontext (en + zh)
Deploy / build (push) Successful in 15s

This commit is contained in:
2026-09-11 12:19:16 +08:00
parent 2645f778b2
commit 30a1963bea
6 changed files with 462 additions and 0 deletions
Binary file not shown.

After

Width:  |  Height:  |  Size: 61 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 44 KiB

+17
View File
@@ -267,6 +267,23 @@ const BANNERS = {
{ n: '5', label: 'verify ✓' },
],
},
'ai-furniture-compositing-with-flux-kontext': {
titlebar: 'client — furniture compositing PoC',
lines: [
{ t: 'prompt', text: '$' }, { t: 'cmd', text: 'POST fal-ai/flux-pro/kontext/multi · 2 photos in' },
{ t: 'err', text: 'toDataURL: tainted canvas · may not be exported' },
{ t: 'prompt', text: '$' }, { t: 'cmd', text: 'fix = same-origin paths · enhance_prompt:false · name objects' },
{ t: 'prompt', text: '' }, { t: 'ok', text: '→ 1 staged room out ✓ ($0.04, ~17s)' },
],
flow: [
{ n: '1', label: '2 photos in' },
{ n: '2', label: 'tainted canvas', err: true },
{ n: '3', label: 'same-origin fix' },
{ n: '4', label: 'prompt anchoring' },
{ n: '5', label: '1 room out ✓' },
],
},
};
const DEFAULT_BANNER = {
+5
View File
@@ -93,6 +93,11 @@ const TERMINALS = {
<div class="line"><span class="prompt">$</span><span class="cmd">git pull --ff-only origin main</span></div>
<div class="line"><span class="prompt">&nbsp;</span><span class="err">refusing: local divergence · never --force</span></div>
<div class="line"><span class="prompt">$</span><span class="cmd">diff -rq live skills · merge on conflict</span><span class="fix">→ synced ✓</span></div>`,
'ai-furniture-compositing-with-flux-kontext': `
<div class="line"><span class="prompt">$</span><span class="cmd">fetch fal-ai/flux-pro/kontext/multi · 2 photos in</span></div>
<div class="line"><span class="prompt">&nbsp;</span><span class="err">toDataURL: tainted canvas · may not be exported</span></div>
<div class="line"><span class="prompt">$</span><span class="cmd">enhance_prompt:false · same-origin paths</span><span class="fix">→ 1 room out ✓</span></div>`,
};
const DEFAULT_TERMINAL = `
@@ -0,0 +1,262 @@
---
title: "How I Built an AI Furniture Compositing Demo With FLUX.1 Kontext"
description: "Compositing furniture photos into styled room scenes with fal.ai FLUX.1 Kontext — the multi-image endpoint, the tainted-canvas fix, and why prompt enhancement had to go."
pubDate: 2026-09-11
category: case-studies
tags: [fal-ai, flux, ai-image, javascript, canvas, ecommerce]
ogImage: /og/ai-furniture-compositing-with-flux-kontext.png
banner: /banners/ai-furniture-compositing-with-flux-kontext.png
---
A furniture client came to me with a problem that every small e-commerce seller
eventually hits: their catalogue photos are **isolated product shots** — a
table on white, a chair on white — but customers don't buy furniture from a
white void. They buy the *room*. Staging a real photoshoot for every product,
in every interior style, is out of the question for a one-person business.
The ask: take two product photos (a dining table, a chair) and place them
together into a believable, high-end interior scene — generated, not
photographed. This post is the story of the proof-of-concept I built to prove
that's possible, the model I chose, and the three bugs that tried to eat it.
## Why FLUX.1 Kontext — and the endpoint that matters
The obvious first instinct is a general text-to-image model: *"a dining table
and a chair in a modern living room."* That fails the moment the client says
*"no — MY table, MY chair, the exact one on my product page."* Generating a
lookalike product is worthless; the whole point is to keep the **real product
unchanged** and only change the room around it.
That's the specific problem **FLUX.1 Kontext** (via [fal.ai](https://fal.ai))
is built for: image-conditioned generation, where the reference image *anchors*
the product and the prompt describes the scene. But there's a subtlety that cost
me an afternoon: there are **two** endpoints.
| Endpoint | Reference input | Use case |
|---|---|---|
| `fal-ai/flux-pro/kontext` | `image_url` (single) | edit / re-context one image |
| `fal-ai/flux-pro/kontext/multi` | `image_urls` (array) | **combine multiple reference images** |
I needed to *fuse two separate products into one scene*, so plain `kontext`
wasn't enough — it only takes one reference image. The `multi` variant accepts
an array of reference images and is what makes "table + chair → one room"
possible. It's flagged experimental, but it's the only game in town for this.
## Calling fal.ai from a plain HTML file — no SDK, no server
For a proof-of-concept I didn't want to stand up a backend. fal.ai exposes a
plain REST queue API, so a single `index.html` with vanilla JavaScript can do
the whole job. There are three steps, not one:
```js
// 1. Submit — returns request_id + polling URLs
const submit = await fetch(
"https://queue.fal.run/fal-ai/flux-pro/kontext/multi",
{
method: "POST",
headers: {
Authorization: "Key " + FAL_KEY,
"Content-Type": "application/json",
},
body: JSON.stringify({
prompt: "...",
image_urls: [tableDataUrl, chairDataUrl],
enhance_prompt: false,
}),
}
);
const { request_id, status_url, response_url } = await submit.json();
// 2. Poll status until COMPLETED (takes ~530s)
let status;
do {
await new Promise((r) => setTimeout(r, 2000));
status = (await (await fetch(status_url, {
headers: { Authorization: "Key " + FAL_KEY },
})).json()).status;
} while (status === "IN_QUEUE" || status === "IN_PROGRESS");
// 3. Fetch the result
const result = await (await fetch(response_url, {
headers: { Authorization: "Key " + FAL_KEY },
})).json();
// result.images[0].url is the generated image
```
Two things I verified early, because they make or break the "single file"
approach:
1. **CORS is open.** `queue.fal.run` returns a permissive
`access-control-allow-origin` and allows the `authorization` header, so a
browser can call it directly with no proxy. I confirmed this with a
preflight before writing any UI.
2. **Local images go in as base64 data URIs.** fal.ai accepts `data:` URIs in
`image_urls`, so I never had to upload the client's photos to a storage
bucket first — the demo reads a file, compresses it, and ships it straight
in the request body.
## The canvas compression layer
Phone photos are 510 MB. Two of those, base64-encoded, bloats the request to
unusable size and slows generation. So before submitting, the demo runs each
image through a `<canvas>` to resize and re-encode it:
```js
async function imgToDataURI(file, maxSize = 1024, quality = 0.85) {
const img = await createImageBitmap(file); // or new Image()
const scale = Math.min(1, maxSize / Math.max(img.width, img.height));
const c = document.createElement("canvas");
c.width = Math.round(img.width * scale);
c.height = Math.round(img.height * scale);
c.getContext("2d").drawImage(img, 0, 0, c.width, c.height);
return c.toDataURL("image/jpeg", quality); // ~200400 KB each
}
```
A 7 MB phone shot becomes a ~300 KB JPEG. Fast to upload, fast for the model to
process.
## Bug #1 — "Tainted canvases may not be exported"
The demo worked in my head. Then it didn't work in a browser. The moment I hit
"generate" I got:
> `Failed to execute 'toDataURL' on 'HTMLCanvasElement': Tainted canvases may not be exported.`
This is a browser security boundary, not a fal.ai problem. When you open the
page over `file://` and load a *local* image with `<img src="chair.jpg">`,
that image is an **opaque origin**. The moment it's drawn onto a canvas, the
canvas becomes *tainted*, and `toDataURL()` refuses to export it — the browser
won't let a webpage read back the pixels of a file it can't prove it's allowed
to read.
Two fixes, and I used both at different points:
1. **Inline the default images as base64 data URIs** directly in the HTML
source. Data URIs don't taint the canvas, so the demo works even when
double-clicked from disk.
2. **Serve over HTTP with same-origin relative paths** (`src="chair.jpg"`).
Once the page and images share an origin, the canvas stays clean. This is
the right answer for the real deployment — the demo now lives at a URL, not
a double-clicked file.
The general rule: **`toDataURL` fails the instant any opaque-origin image touches
the canvas.** If you're loading local files, either inline them or serve over
HTTP. There is no third way without changing browser security settings.
## Bug #2 — the API key is in the HTML
For a demo this is tolerable; for anything public it's a hole. A `Key` header in
a client-side `fetch` means the key is in the page source, readable by anyone
who presses F12. Fine for a PoC I hand to a client, unacceptable for production.
The plan for the real WordPress plugin (the next phase) is to move the key to
**server-side PHP**: the plugin's endpoint does the authenticated fal.ai call
and returns the image, while the browser only talks to the plugin's own route.
The demo proved the pipeline works; the plugin will put the secret where
secrets belong.
## Bug #3 — the output looked like "the second image"
This was the one that made me question whether `multi` even worked. I swapped in
a new table photo, generated, and the result looked almost identical to the
chair reference — as if the model had just ignored the table and copied the
chair.
Before blaming the model, I checked something falsifiable: **is `multi` even
compositing, or just echoing one input?** I compared the generated image
against both source images with a perceptual hash and mean pixel difference:
| Comparison | Mean pixel diff (0 = identical) | dHash distance |
|---|---|---|
| result vs table | 55.7 | 31 |
| result vs chair | 55.3 | 29 |
| table vs chair | 69.3 | 30 |
The result was *far* from both inputs — the endpoint genuinely composites a new
scene, not a copy. So the "looks like the chair" problem wasn't the API; it was
the **inputs and the prompt**.
Three compounding causes:
1. **The reference photos were inconsistent.** The original `chair.jpg` was a
photo of a full table-and-chair *scene*, not a clean single chair — so the
model already saw "table + chair" satisfied by one image and leaned on it.
2. **Positional prompting is unreliable.** Kontext matches images by *content*,
not array order. Saying "the first image" / "the second image" doesn't
reliably map to "the table" / "the chair". The fix is to name the objects —
*"the dining table"* and *"the chair"* — and let the model match them to the
right reference.
3. **Prompt enhancement was on.** fal.ai's `enhance_prompt` (default off, but I
had it in mind) rewrites your prompt into a richer aesthetic description —
exactly wrong when the goal is *"do not reinterpret the product."* I pinned
it to `false` so the model obeys the literal prompt instead of embellishing
it.
## The prompt that finally worked
The client's core requirement was strict: **preserve the furniture exactly** and
only invent the room. That needs a prompt that spends most of its length
*preventing* reinterpretation, not describing style:
> Create a photorealistic interior photograph using the exact dining table and
> exact chair from the reference images. Preserve both furniture pieces exactly
> as shown — do not redesign, recolor, repaint, restyle, replace, or reinterpret
> either product. Keep their original geometry, proportions, construction,
> material, finish, texture, and color exactly unchanged.
>
> Place the chair naturally beside the dining table in a realistic dining
> position, slightly pulled under the table and properly aligned with it.
> Maintain realistic scale, perspective, and physical contact with the floor.
> The chair and table must look photographed together in the same real room,
> with natural contact shadows — not composited or pasted together.
>
> Preserve the true original color. Do not allow the room lighting, wall or
> floor colors, or grading to alter the furniture.
>
> Set the scene in a spacious modern living room with floor-to-ceiling windows
> and subtle warm afternoon daylight. Use restrained, neutral interior colors.
>
> Photorealistic commercial furniture photography, realistic camera
> perspective, natural proportions, high detail.
The demo ships six of these, one per scene style (modern living room, cozy
dining room, Scandinavian, wabi-sabi, and a clean studio), each differing only
in the "set the scene" paragraph.
## The result
One `index.html`, no backend, no SDK — two product photos in, a staged room
out, at **$0.04 per image** and roughly 1520 seconds of generation. The client
now has a working preview they can show *their* customers, with a scene selector
and an editable prompt, before committing to the full WordPress plugin.
The proof-of-concept answered the question it was built to answer: **yes, you
can keep the real product and only generate the room around it** — as long as
you respect the three constraints that actually matter: a clean reference image
per product, object-named (not positional) prompts, and `enhance_prompt: false`.
## What I'd do differently
- **Require clean, single-product reference images on day one.** Every failure
mode downstream — the "looks like one image" problem, the proportion drift —
traces back to inconsistent source photos. I'd ship a tiny client-facing
guide ("one product per photo, plain background, no other furniture") before
touching any model.
- **Turn `enhance_prompt` off explicitly, first.** I lost a generation to the
enhancement rewriting my careful "do not reinterpret" instructions into
exactly the opposite.
- **Validate CORS with a real preflight before building UI.** It's a two-line
`curl` and it de-risks the entire architecture decision.
---
## Want this for your product catalogue?
I build AI image pipelines, WordPress plugins, and self-hosted infrastructure
for e-commerce businesses. If you want to stage your products in styled room
scenes — or hire me for a similar AI integration — I'd love to talk:
- 📱 **WhatsApp:** [+60 12-797 2969](https://wa.me/60127972969)
- 📧 **Email:** [[email protected]](mailto:[email protected]?subject=AI%20product%20image%20compositing)
- 🌐 **Website:** [hoelee.com](https://hoelee.com)
@@ -0,0 +1,178 @@
---
title: "我如何用 FLUX.1 Kontext 打造 AI 家具合成演示"
description: "用 fal.ai FLUX.1 Kontext 把家具产品照合成到风格化房间场景——多图端点、tainted canvas 的修复,以及为什么必须关掉 prompt enhancement。"
pubDate: 2026-09-11
category: case-studies
tags: [fal-ai, flux, ai-image, javascript, canvas, ecommerce]
ogImage: /og/ai-furniture-compositing-with-flux-kontext.png
banner: /banners/ai-furniture-compositing-with-flux-kontext.png
---
一位家具客户带着一个每个小电商卖家最终都会撞上的问题来找我:他们的产品图都是**孤立的产品照**——一张白底桌子、一张白底椅子——但顾客不会从一片白茫茫的虚无里买家具,他们买的是*房间*。给每个产品、每种室内风格都拍一次实体棚拍,对一个人的生意来说根本不可能。
需求是:拿两张产品照(一张餐桌、一把椅子),把它们一起放进一个可信的、高级感的室内场景——生成出来的,而不是拍出来的。这篇文章讲的就是我为验证这件事可行而做的概念验证:我选的模型,以及三个差点把它搞砸的 bug。
## 为什么选 FLUX.1 Kontext——以及那个关键的端点
第一反应自然是通用文生图模型:*「一张餐桌和一把椅子放在现代客厅里」*。可当客户说*「不——是我的桌子、我的椅子,就是我产品页上那一款」*时,这个思路就崩了。生成一个相似品毫无价值;重点是要让**真实产品原封不动**,只改变它周围的房间。
这正是 **FLUX.1 Kontext**(通过 [fal.ai](https://fal.ai))专门解决的问题:以图生图,参考图*锚定*产品,提示词描述场景。但有个坑让我耗掉一个下午:这里有**两个**端点。
| 端点 | 参考图输入 | 用途 |
|---|---|---|
| `fal-ai/flux-pro/kontext` | `image_url`(单张) | 编辑 / 重新语境化一张图 |
| `fal-ai/flux-pro/kontext/multi` | `image_urls`(数组) | **合并多张参考图** |
我需要把*两个独立产品融合进一个场景*,所以普通 `kontext` 不够——它只吃一张参考图。`multi` 变体接受参考图数组,这才是「桌子 + 椅子 → 一个房间」能成立的关键。它被标记为实验版,但这件事上它是唯一的选择。
## 用一个纯 HTML 文件调用 fal.ai——不用 SDK,不用服务器
做概念验证我不想搭后端。fal.ai 暴露了一个纯 REST 队列 API,所以一个 `index.html` 配 vanilla JavaScript 就能搞定全部。但这是三步,不是一步:
```js
// 1. 提交——返回 request_id + 轮询 URL
const submit = await fetch(
"https://queue.fal.run/fal-ai/flux-pro/kontext/multi",
{
method: "POST",
headers: {
Authorization: "Key " + FAL_KEY,
"Content-Type": "application/json",
},
body: JSON.stringify({
prompt: "...",
image_urls: [tableDataUrl, chairDataUrl],
enhance_prompt: false,
}),
}
);
const { request_id, status_url, response_url } = await submit.json();
// 2. 轮询状态直到 COMPLETED(约 530 秒)
let status;
do {
await new Promise((r) => setTimeout(r, 2000));
status = (await (await fetch(status_url, {
headers: { Authorization: "Key " + FAL_KEY },
})).json()).status;
} while (status === "IN_QUEUE" || status === "IN_PROGRESS");
// 3. 取结果
const result = await (await fetch(response_url, {
headers: { Authorization: "Key " + FAL_KEY },
})).json();
// result.images[0].url 就是生成的图片
```
有两件事我很早就验证了,因为它们决定了「单文件」方案成不成立:
1. **CORS 是开放的。** `queue.fal.run` 返回宽松的 `access-control-allow-origin`,并允许 `authorization` 头,所以浏览器可以直接调用,无需代理。我在写任何 UI 之前就用 preflight 确认了这一点。
2. **本地图以 base64 data URI 传入。** fal.ai 的 `image_urls` 接受 `data:` URI,所以我根本不需要先把客户的照片传到存储桶——demo 直接读文件、压缩,然后塞进请求体。
## Canvas 压缩层
手机照片有 5–10 MB。两张这样、再 base64 编码,请求体会膨胀到无法使用,生成也会变慢。所以提交前,demo 会把每张图过一遍 `<canvas>` 来缩放并重新编码:
```js
async function imgToDataURI(file, maxSize = 1024, quality = 0.85) {
const img = await createImageBitmap(file); // 或 new Image()
const scale = Math.min(1, maxSize / Math.max(img.width, img.height));
const c = document.createElement("canvas");
c.width = Math.round(img.width * scale);
c.height = Math.round(img.height * scale);
c.getContext("2d").drawImage(img, 0, 0, c.width, c.height);
return c.toDataURL("image/jpeg", quality); // 每张约 200400 KB
}
```
一张 7 MB 的手机照变成约 300 KB 的 JPEG。上传快,模型处理也快。
## Bug #1 ——「Tainted canvases may not be exported」
demo 在我脑子里能跑,在浏览器里却不行。我一点「生成」就收到:
> `Failed to execute 'toDataURL' on 'HTMLCanvasElement': Tainted canvases may not be exported.`
这是浏览器的安全边界,不是 fal.ai 的问题。当你用 `file://` 打开页面、并用 `<img src="chair.jpg">` 加载一张*本地*图片时,那张图属于**不透明源**。它一被画到 canvas 上,canvas 就变成 *tainted*`toDataURL()` 拒绝导出——浏览器不会让网页读回一张它无法证明有权读取的文件的像素。
有两个修复,我在不同阶段都用过:
1. **把默认图以 base64 data URI 内嵌进 HTML 源码。** data URI 不会污染 canvas,所以即使从磁盘双击打开也能跑。
2. **走 HTTP、用同源相对路径**`src="chair.jpg"`)。一旦页面和图片同源,canvas 就保持干净。这是真实部署的正确答案——demo 现在跑在一个 URL 上,而不是双击打开的文件。
通用规则:**只要任何不透明源的图片碰到 canvas,`toDataURL` 立刻报错。** 如果你要加载本地文件,要么内嵌,要么走 HTTP。不改浏览器安全设置,没有第三条路。
## Bug #2 —— API key 就在 HTML 里
对 demo 来说这可以容忍;对任何公开的东西来说这是个洞。客户端 `fetch` 里的 `Key` 头意味着 key 就在页面源码里,任何人按 F12 就能读到。给我交给客户的概念验证没问题,上生产就不可接受。
真正 WordPress 插件(下一阶段)的计划是把 key 挪到**服务端 PHP**:插件的端点去做带鉴权的 fal.ai 调用并返回图片,浏览器只跟插件自己的路由对话。demo 证明了管线能跑通;插件会把秘密放回秘密该待的地方。
## Bug #3 —— 输出看起来「像第二张图」
这是让我怀疑 `multi` 到底能不能用的那个。我换了一张新桌子图、生成,结果看起来跟椅子参考图几乎一模一样——好像模型直接忽略了桌子、复制了椅子。
在怪模型之前,我先查了一件可证伪的事:**`multi` 到底是在合成,还是只是在回显某一张输入?** 我用感知哈希和平均像素差对比了生成图和两张源图:
| 对比 | 平均像素差(0 = 相同) | dHash 距离 |
|---|---|---|
| 结果 vs 桌子 | 55.7 | 31 |
| 结果 vs 椅子 | 55.3 | 29 |
| 桌子 vs 椅子 | 69.3 | 30 |
结果离两张输入都*很远*——端点是真在合成新场景,不是复制。所以「像椅子」的问题不在 API,而在**输入和提示词**。
三个叠加的原因:
1. **参考图不一致。** 最初的 `chair.jpg` 是一张完整的「桌+椅」*场景*照,不是干净的单椅——于是模型看到一张图就已经满足「桌+椅同框」,就倾向照着它来。
2. **按位置指代不可靠。** Kontext 按图片*内容*匹配,而不是数组顺序。说「第一张图」/「第二张图」并不能可靠地映射到「桌子」/「椅子」。修复是直接*命名物体*——*「the dining table」*和*「the chair」*——让模型自己去匹配正确的参考图。
3. **prompt enhancement 开着。** fal.ai 的 `enhance_prompt`(默认关,但我一直惦记着它)会把你的提示词改写成更丰富的审美描述——当目标是*「别重新解读产品」*时,这正好帮倒忙。我把它钉死为 `false`,让模型遵守字面提示词,而不是添油加醋。
## 最终起作用的提示词
客户的核心要求很严格:**原样保留家具**,只去发明房间。这需要一段把大部分篇幅花在*阻止*重新解读、而不是描述风格上的提示词:
> Create a photorealistic interior photograph using the exact dining table and
> exact chair from the reference images. Preserve both furniture pieces exactly
> as shown — do not redesign, recolor, repaint, restyle, replace, or reinterpret
> either product. Keep their original geometry, proportions, construction,
> material, finish, texture, and color exactly unchanged.
>
> Place the chair naturally beside the dining table in a realistic dining
> position, slightly pulled under the table and properly aligned with it.
> Maintain realistic scale, perspective, and physical contact with the floor.
> The chair and table must look photographed together in the same real room,
> with natural contact shadows — not composited or pasted together.
>
> Preserve the true original color. Do not allow the room lighting, wall or
> floor colors, or grading to alter the furniture.
>
> Set the scene in a spacious modern living room with floor-to-ceiling windows
> and subtle warm afternoon daylight. Use restrained, neutral interior colors.
>
> Photorealistic commercial furniture photography, realistic camera
> perspective, natural proportions, high detail.
demo 里内置了六段这样的提示词,每种场景风格一段(现代客厅、温馨餐厅、北欧风、侘寂风、以及干净的影棚),每一段只改「set the scene」那一节。
## 结果
一个 `index.html`,没有后端、没有 SDK——两张产品照进,一个布置好的房间出,**每张图 $0.04**,生成约 15–20 秒。客户现在有了一个能拿给*他们*顾客看的可用预览,带场景选择器和可编辑提示词,然后再决定是否投入完整的 WordPress 插件。
这个概念验证回答了一个它生来就要回答的问题:**是的,你可以保留真实产品、只生成它周围的房间**——只要你尊重那三条真正要紧的约束:每个产品一张干净的参考图、按物体命名(而非按位置)的提示词、以及 `enhance_prompt: false`
## 我会做不同的地方
- **第一天就要求干净的单品参考图。** 下游的每一种失败模式——「像某一张图」的问题、比例漂移——都能追溯到不一致的源照片。我会在碰任何模型之前,先发一份给客户的简短说明(「每张照片一个产品、纯色背景、没有其他家具」)。
- **先显式关掉 `enhance_prompt`。** 我浪费了一次生成,就因为 enhancement 把我精心写的「别重新解读」改写成恰恰相反的东西。
- **在写 UI 之前先用真实 preflight 验证 CORS。** 两行 `curl` 的事,却能给整个架构决策祛除风险。
---
## 想为你的产品目录做这个吗?
我为电商企业构建 AI 图像管线、WordPress 插件,以及自托管基础设施。如果你想把产品摆进风格化的房间场景——或者想雇我做类似的 AI 集成——我非常乐意交流:
- 📱 **WhatsApp** [+60 12-797 2969](https://wa.me/60127972969)
- 📧 **邮箱:** [[email protected]](mailto:[email protected]?subject=AI%20product%20image%20compositing)
- 🌐 **网站:** [hoelee.com](https://hoelee.com)