Quick Start
FigCraft is a desktop app for AI image & video creation. At its core is an agent that takes action — describe what you want in plain language and it generates, edits, cuts out, makes videos, batches outputs, and runs multi-step tasks for you.
Install
- 1Download the installer for your OS (macOS .dmg / Windows .exe).
- 2macOS: drag into Applications. Windows: double-click to install.
- 3On first launch, if the OS flags an unknown source, allow it in system settings.
- 4The app checks for updates automatically and offers one-click download.
Sign in / Sign up
- 1Sign in with a phone number + SMS code; a new account is created automatically if needed.
- 2Have an invite code? Enter it in the optional "Invite code" field to be linked to the inviter.
- 3Google sign-in is also supported.
Your first image
- 1Open the Image page and describe the scene in the input box.
- 2Need consistency? Drag or paste reference images (same character/product).
- 3Press Enter — the agent generates onto the canvas; keep chatting to refine.
The Agent
The FigCraft agent does more than make one image. It understands a goal, plans steps, and calls tools to finish the whole job. Say "make 5 hero shots of these shoes in different scenes" and it analyzes, generates, compares, edits if needed, and delivers.
Consistency tip: add a character/product photo as a reference — the agent keeps the same identity across a whole set.
How to Use
Chat is the control
Everything works through chat. Just state the goal: "swap to a clean white background", "upscale to 4K", "make 3 expressions of this character".
References & consistency
- •Drag or paste images as references — multiple at once (character, product, style).
- •To keep the same person/product, say "keep the character from the reference".
- •For local edits say "only change the XX area" — it inpaints instead of redrawing.
Canvas & library
- •Results live on the canvas; keep editing or screenshot-analyze from it.
- •Save frequent products/characters/scenes to the library and reuse in one line.
- •History keeps past sessions/tasks; long tasks resume from checkpoint.
Video
- •Text-to-video, image-to-video, first/last-frame; chain clips via the last frame for longer videos.
- •Some models include audio and multilingual lip-sync.
Billing
Everything is billed in credits, from two sources: a monthly plan allowance (resets monthly) and top-ups (never expire). Plan allowance is used first, then permanent credits.
What a plan includes
- •Plan credits — for image / video generation.
- •Advanced chat quota — for premium models (GPT / Claude / Gemini Pro).
- •Basic chat quota — for basic models (Qwen / DeepSeek).
How chat is charged
Within a plan, chat is counted as requests — simple and predictable. Pricier models cost more requests per turn (e.g. GPT-5.5 = 1, Claude Opus = 3). Once used up, chat switches to usage-based (tokens) against permanent credits.
How image / video is charged
Per model: images by credits-per-image, video by credits-per-second. Stronger models cost more. Plan credits first, then permanent. Exact rates are shown in the app.
Reference: 1 credit ≈ $0.01 ≈ ¥0.072 ($1 = 100 credits). Permanent credits never expire; plan credits reset monthly.
Agent Tools
The agent has a full toolset across 13 categories, gated by a three-tier risk-based permission model and invoked automatically — you rarely need the details, just a sense of what it can do.
13 Categories
What It Can Do
Models
The app aggregates leading image & video models. The agent picks one automatically, and you can specify one manually.
Image models · 7
Video models · 10
Quick picks: transparent background → GPT Image 2; character consistency → Nano Banana Pro; 4K video → Veo 3.1; long video → Seedance last-frame chaining.
Architecture
The engineering behind the toolset — the 10 things that make the agent a production-grade system, not a gimmick.
Closed-Loop Reasoning
Most AI agents on the market are pipelines — a fixed sequence of LLM calls that freezes the moment reality deviates from the script.
FigCraft's image agent runs true closed-loop reasoning: every iteration re-inspects the canvas, reference pool, and decision history, then dynamically picks the next tool to invoke — up to 200 iterations per task.
- •Every turn the LLM re-evaluates from scratch — no scripted path
- •Tool results feed directly into the next round's decision
- •A single tool failure never crashes the task — the agent diagnoses the error and adapts its strategy
Three-Tier Tool Permissions
All 61 agent tools are tiered by blast radius, giving brands the confidence to delegate real authority to AI.
- •Read-only tools (analyze, search, capture) — executed in parallel for maximum throughput
- •Mutating tools (generate, edit, composite, export) — serialized to prevent concurrent conflicts
- •High-sensitivity tools (shell, overwrite, bulk delete) — every invocation requires explicit user confirmation, no agent bypass
- •Tools can emit a terminal signal to halt the loop immediately, preventing wasted tokens
Multi-SKU Subject Consistency
The hardest problem in apparel: shoot one jacket across 30 colorways and it looks like 30 different models wearing it — random latent drift fragments the subject.
FigCraft is purpose-built for ecommerce subject consistency — however large the set, the subject stays locked.
- •Auto-detects the task type and picks the right generation mode for singles, grouped variants, or sequential evolution
- •Can keep every frame faithful to the original product or person you uploaded
- •Or lock the whole set to the look of a single hero frame for maximum style uniformity
- •Supports carrying a prior frame's overall look forward across the set
- •Consistency choices are written into the execution plan — visible and editable before approval
Zero Surprise Spend
Before any multi-step task, the agent surfaces the full plan: "Generate 1 white-background hero + 3 alpine scenes + 2 desert scenes, estimated 12 credits, each anchored to user upload."
Three choices for the user: approve, cancel, or refine in natural language. No credits are spent until approval.
- •Zero credits consumed before approval — no generation calls during planning
- •Iterative revision — refine the plan as many times as needed
- •Approved plans are archived for full credit traceability after the fact
Intelligent Result Cache
Mid-tier LLMs share a familiar flaw: they invoke the same tool twice, sometimes three times, burning tokens each round.
We reuse results within a single run — repeated read-only operations never burn compute twice; the prior result is reused instantly.
- •Repeated read-only operations auto-reuse their result within a run
- •Duplicate calls are recognized precisely, avoiding wasted re-execution
- •Complex tasks see a significant reduction in token spend
Persistent Long-Session Memory
A single apparel shoot can produce hundreds of images across dozens of conversation turns. Most frameworks either blow the context window or start hallucinating.
Two memory layers: short-term auto-summarization, and long-term persistence of the essentials.
- •Short-term — as a conversation grows, early turns are auto-summarized to free up context
- •Long-term — key information is remembered persistently and carried across sessions
- •Every image carries its provenance (who uploaded it, which turn, what was asked) — no mix-ups
- •At turn 80, the agent still remembers what was uploaded at turn 3
Crash-Resilient Task System
In enterprise environments, crashes, power loss, and forced restarts are routine. Most AI tools lose the in-flight task entirely.
Our task system uses a three-layer design that recovers from crashes with zero data loss.
- •Task IDs are incremental strings (1 / 2 / 3) rather than UUIDs — lighter cognitive load for the model, more reliable dispatch
- •Subtasks use positional indices (0 / 1 / 2) so the model never has to memorize long strings
- •Up to 100 tasks are persisted locally for inspection and recovery at any time
- •Auto-repair on restart — orphaned "in-progress" tasks are demoted to paused, eliminating ghost jobs
Network Fault Tolerance
Anyone shipping against third-party LLM APIs knows the drill — occasional timeouts, 500s, and rate limits. The agent has to absorb all of it.
- •120s API timeout, generous enough for extended-thinking models
- •3 retries with exponential backoff (500ms → 1s → 2s)
- •4xx client errors fail fast; 5xx, 429, and timeouts auto-retry
- •Malformed responses are treated as failures and retried rather than passed through as empty
- •Empty responses surface explicit errors (likely safety filtering or max_tokens exhausted by thinking) instead of silent exits
Dynamic System Prompt
Most agents bake the system prompt once at startup — they never see that the canvas changed, the model swapped, or new references were uploaded.
Our agent rebuilds the system prompt every iteration, injecting live canvas state, reference pool, current selection, and model capabilities.
- •Real-time awareness of canvas contents, aspect ratio, and resolution
- •Reference pool overview is indexed per-slot so the agent never mixes them up
- •Available models and current-model capabilities (multi-reference, inpainting, max-N) are injected live
- •Switch models and the agent immediately knows the new capability surface
Full Event Transparency
What the agent is thinking, which tool it invoked, what came back, why it's requesting confirmation — every signal streams to the UI in real time.
Clients see every decision the agent makes — a completely different trust profile from the "spinner-until-result" black-box tools.
- •Event types — thinking, tool_call, tool_result, message, permission_request, error
- •Tool-call arguments are exposed live, letting clients reverse-engineer the agent's reasoning
- •Errors are humanized into actionable advice — switch model, simplify request, and so on