TECHNOLOGY
One sentence, a whole visual system
Once engineers had a coding agent, they stopped typing line by line. Images deserve the same thing — an agent that actually does the work on your own machine: it lays out the infinite canvas itself, reads the assets on your disk, and takes direction from your phone or WeChat when you are away from the desk. This page is how it works.
The agent lays out the board itself
You do not place nodes one by one. Say what you want; it creates the nodes, wires them up, fires off a batch, and arranges the next one when results land.
Live demo — the agent creating nodes, wiring them and generating on the canvas
It creates and wires the nodes
Upstream output feeds the next node, building a multi-step chain. You watch it think and build, instead of staring at a progress bar that dumps images at the end.
Fires a batch, waits for results
A whole group goes out at once and each lands as it returns. Only when the batch is complete does it decide what to do next.
It keeps the board tidy
Node placement, whether a parameter panel is open, what the notes on the canvas say — none of it is left for you to clean up.
Spending waits for your approval
Anything that costs credits goes through an approval bar first: what this batch will make and what it will cost, shown before it runs.
Put it to work while you are away from the desk
Two channels, both set up by scanning a code. The point is not that it works remotely — it is which path a remote instruction takes.
Pair by scanning a code on the same WiFi and your phone becomes a remote keyboard for this machine. The link carries a one-time token and only listens on the local network — connections without the token are refused.
phone -> local network -> your machine
Scan to bind WeChat to this machine, then just say what you want in a chat — voice messages work too. It runs Tencent’s official ClawBot plugin, installed on your own machine.
WeChat -> Tencent -> your machine
This is the whole point: a remote instruction does not start a second, server-side agent. It is typed into the desktop input box and sent. From there it follows exactly the same flow as a sentence you typed yourself.
A server-side design would have to drop all four — the server has neither your canvas nor your disk. And messages on both channels stay off FigCraft servers entirely.
Anyone can wire up a model. Finishing the job is the hard part
Nano Banana, Seedream and Veo are open to everyone — plugging them in is not a moat. Turning one sentence into a finished set, without drifting or stalling halfway, is the work our own agent does.
It plans first, you approve, then it runs
Anything multi-step starts as a plan: how many pieces, which models, which brand rules. Nothing is spent until you approve it.
Plan, approve, execute — three separate steps
Long jobs do not lose the thread
Dozens of steps in, it still holds the brand rules you set at the start. Oversized intermediate results are compressed in place, and the whole run is summarized as it nears the limit.
100k-token context budget per job, auto-compressed and retried when it overflows
It routes around a flaky model
When an upstream model times out, rate-limits or drops out, it switches within the same turn and keeps going instead of handing you an error.
Falls back to a backup model once retries run out, inside the same turn
Every step is in front of you
Which file it read, which model it called, what it produced, where it got stuck — all listed as it happens. Stop it mid-run if it is going the wrong way.
Tool calls, arguments and results all visible
A loop, not a pipeline
It looks at what it just produced and redoes the parts that are wrong, instead of carrying a bad result all the way to delivery.
Up to 200 self-correcting turns per job
It runs on your own machine
Reads project assets and brand rules straight from disk, and writes finals back to the folder you choose. No upload-then-download round trip.
Reads and writes local disk; finals land in ~/Downloads or a folder you pick
The surface it works on
Generation is one part of it. The agent reaches for each of these on its own, and chains them together.
Generate and edit
Make images, revise them, repaint a selected area, cut out subjects, upscale.
Read the picture
Understands subject, light, composition and style, and decides the next edit from that.
Read and write your files
Walk folders, read your brand doc, batch-process assets, write finals back where you asked.
Drive the infinite canvas
Create nodes, wire them, generate in batches, rearrange the board, update its notes.
Break down and plan
Splits one sentence into subtasks, marks what blocks what, and revises the plan mid-run.
Search the web
Finds references, checks specs, looks at how competitors shoot it — then starts work.
Remember preferences
Holds on to your brand rules and revision habits so you do not repeat yourself.
Schedule and automate
Save a flow that worked, then re-run it on a schedule or against a new batch of assets.
Models are a replaceable part
Which is why they come last on this page.
Nano Banana Pro, Seedream 5.0, GPT-5.4 Image 2, Veo 3.1, Seedance 2.5 — every one of them is open to everyone, and wiring them up is not an advantage. The agent picks per subtask, or you name one yourself; when a model is unavailable it switches and keeps going. The model pool changes every month. The engineering above does not.