FigCraftFigCraft
GETTING STARTED

Quick Start

FigCraft is a desktop app for AI image & video creation. At its core is an agent that takes action — describe what you want in plain language and it generates, edits, cuts out, makes videos, batches outputs, and runs multi-step tasks for you.

Install

  1. 1Download the installer for your OS (macOS .dmg / Windows .exe).
  2. 2macOS: drag into Applications. Windows: double-click to install.
  3. 3On first launch, if the OS flags an unknown source, allow it in system settings.
  4. 4The app checks for updates automatically and offers one-click download.

Sign in / Sign up

  1. 1Sign in with a phone number + SMS code; a new account is created automatically if needed.
  2. 2Have an invite code? Enter it in the optional "Invite code" field to be linked to the inviter.
  3. 3Google sign-in is also supported.

Your first image

  1. 1Open the Image page and describe the scene in the input box.
  2. 2Need consistency? Drag or paste reference images (same character/product).
  3. 3Press Enter — the agent generates onto the canvas; keep chatting to refine.
CORE CONCEPTS

The Agent

The FigCraft agent does more than make one image. It understands a goal, plans steps, and calls tools to finish the whole job. Say "make 5 hero shots of these shoes in different scenes" and it analyzes, generates, compares, edits if needed, and delivers.

Generate & edit
Text/image-to-image, inpainting, cutout/background swap, batch, text/image-to-video.
Understand & analyze
Reads uploads, outputs or the live canvas; analyzes products & suggests visuals.
Multi-step tasks
Breaks complex jobs into tasks by dependency; can pause/resume.
Plan mode
Researches first, proposes a plan for approval, then executes.
Files & system
Reads/writes files, imports/exports, runs ffmpeg/sips.
Web
Searches the web and fetches pages for reference.

Consistency tip: add a character/product photo as a reference — the agent keeps the same identity across a whole set.

CORE CONCEPTS

How to Use

Chat is the control

Everything works through chat. Just state the goal: "swap to a clean white background", "upscale to 4K", "make 3 expressions of this character".

References & consistency

  • •Drag or paste images as references — multiple at once (character, product, style).
  • •To keep the same person/product, say "keep the character from the reference".
  • •For local edits say "only change the XX area" — it inpaints instead of redrawing.

Canvas & library

  • •Results live on the canvas; keep editing or screenshot-analyze from it.
  • •Save frequent products/characters/scenes to the library and reuse in one line.
  • •History keeps past sessions/tasks; long tasks resume from checkpoint.

Video

  • •Text-to-video, image-to-video, first/last-frame; chain clips via the last frame for longer videos.
  • •Some models include audio and multilingual lip-sync.
ACCOUNT

Billing

Everything is billed in credits, from two sources: a monthly plan allowance (resets monthly) and top-ups (never expire). Plan allowance is used first, then permanent credits.

What a plan includes

  • •Plan credits — for image / video generation.
  • •Advanced chat quota — for premium models (GPT / Claude / Gemini Pro).
  • •Basic chat quota — for basic models (Qwen / DeepSeek).

How chat is charged

Within a plan, chat is counted as requests — simple and predictable. Pricier models cost more requests per turn (e.g. GPT-5.5 = 1, Claude Opus = 3). Once used up, chat switches to usage-based (tokens) against permanent credits.

How image / video is charged

Per model: images by credits-per-image, video by credits-per-second. Stronger models cost more. Plan credits first, then permanent. Exact rates are shown in the app.

Reference: 1 credit ≈ $0.01 ≈ ¥0.072 ($1 = 100 credits). Permanent credits never expire; plan credits reset monthly.

REFERENCE

Agent Tools

The agent has a full toolset across 13 categories, gated by a three-tier risk-based permission model and invoked automatically — you rarely need the details, just a sense of what it can do.

13 Categories

01
GenerationCore generative capabilities of the agent
4
02
Editing & AnalysisVisual understanding plus AI-driven retouching
5
03
FilesystemDirect local read/write with zero cloud relay
8
04
Shell & I/OOS-level access primitives
3
05
NetworkPull context from the live web
2
06
TasksCrash-resilient task system with zero loss
6
07
User InteractionHigh-sensitivity actions require explicit confirmation
3
08
Multi-AgentAgents that delegate to other agents
7
09
WorkflowPlan Mode with dynamic tool discovery
4
10
Memory & TeamLong-session recall plus team messaging
7
11
Code & ScheduleCron, LSP and worktree primitives
12
12
Asset LibraryAsset catalog + global style anchor for cross-shot consistency
8
13
Video Line & VoiceContinuous shots + per-character voices for long films
8

What It Can Do

Image & Video Generation
Generate a single image (text/image-to-image)
Generate image by image while keeping the character/product consistent
Generate multiple images in one batch
Generate video (text/image, first/last-frame)
Local edits to a specified area with a reference image
Set ratio / quality / model
Video & Voice
Stitch and reorder multiple clips
Text-to-speech voiceover and voice cloning
Understand & Analyze
Vision analysis of uploads, outputs, or the live canvas
Product selling points and visual strategy
Files & System
Read/write local files, import and export (jpg/png/psd)
Invoke local tools for format and media processing
Web
Search the web for references and inspiration
Fetch page content as reference material
Tasks & Planning
Break complex jobs into multi-step tasks, with pause/resume
Plan mode: propose a plan for approval before acting
Memory & Library
Remember brand/character/style across long sessions
Product & asset libraries: organize and reuse in one line
Multi-agent
Spawn sub-agents to handle large jobs in parallel
Agents coordinate and message one another
Interaction
Reply, ask questions, and confirm high-sensitivity actions
REFERENCE

Models

The app aggregates leading image & video models. The agent picks one automatically, and you can specify one manually.

Image models · 7

Nano Banana
Google · conversational editing, SynthID watermark
Nano Banana Pro
Google · up to 14 references, 1K/2K/4K, thinking on
Nano Banana 2 (Flash)
Google · 512px–4K, extreme ratios, adjustable thinking
Seedream 5.0
Volcano · web search, deep thinking, up to 15 per run
Seedream 4.5
Volcano · 2K/4K, up to 15 per run
万相 2.6 (wan2.6-image)
Tongyi · image edit, multi-image fusion, negative prompts
GPT-5.4 Image 2
OpenAI · custom sizes, quality tiers; no transparent bg

Video models · 10

Seedance 2.0
Volcano · web search, last-frame chaining, up to 1080p
Seedance 2.0 Fast
Volcano · fast variant, up to 720p
Seedance 1.5 Pro
Volcano · lip-sync, camera lock, sample mode
Veo 3.1
Google · native 4K, audio, timestamped prompts
万象 I2V Flash (wan2.6)
Tongyi · image-to-video, 720p/1080p
万象 R2V Flash (wan2.6)
Tongyi · image+video references (≤5)
万象 KF2V Flash (wan2.2)
Tongyi · fixed 5s, first/last-frame, effect templates
HappyHorse 1.0
Alibaba · native A/V + 7-language lip-sync, multi-shot
HappyHorse 1.0 图生视频
Alibaba · from first frame, inherits aspect ratio
万相 2.7 文生视频
Tongyi · wan2.6 upgrade, multi-shot narrative

Quick picks: transparent background → GPT Image 2; character consistency → Nano Banana Pro; 4K video → Veo 3.1; long video → Seedance last-frame chaining.

DEEP DIVE

Architecture

The engineering behind the toolset — the 10 things that make the agent a production-grade system, not a gimmick.

01

Closed-Loop Reasoning

Not a pipeline — a 200-iteration closed loop · Loop Reasoning

Most AI agents on the market are pipelines — a fixed sequence of LLM calls that freezes the moment reality deviates from the script.

FigCraft's image agent runs true closed-loop reasoning: every iteration re-inspects the canvas, reference pool, and decision history, then dynamically picks the next tool to invoke — up to 200 iterations per task.

  • •Every turn the LLM re-evaluates from scratch — no scripted path
  • •Tool results feed directly into the next round's decision
  • •A single tool failure never crashes the task — the agent diagnoses the error and adapts its strategy
02

Three-Tier Tool Permissions

Hand over the keys without giving up the wheel · Permission Tiers

All 61 agent tools are tiered by blast radius, giving brands the confidence to delegate real authority to AI.

  • •Read-only tools (analyze, search, capture) — executed in parallel for maximum throughput
  • •Mutating tools (generate, edit, composite, export) — serialized to prevent concurrent conflicts
  • •High-sensitivity tools (shell, overwrite, bulk delete) — every invocation requires explicit user confirmation, no agent bypass
  • •Tools can emit a terminal signal to halt the loop immediately, preventing wasted tokens
03

Multi-SKU Subject Consistency

One set, one subject · Shared-Subject Strategy

The hardest problem in apparel: shoot one jacket across 30 colorways and it looks like 30 different models wearing it — random latent drift fragments the subject.

FigCraft is purpose-built for ecommerce subject consistency — however large the set, the subject stays locked.

  • •Auto-detects the task type and picks the right generation mode for singles, grouped variants, or sequential evolution
  • •Can keep every frame faithful to the original product or person you uploaded
  • •Or lock the whole set to the look of a single hero frame for maximum style uniformity
  • •Supports carrying a prior frame's overall look forward across the set
  • •Consistency choices are written into the execution plan — visible and editable before approval
04

Zero Surprise Spend

Every multi-step task ships a plan first · Plan Approval

Before any multi-step task, the agent surfaces the full plan: "Generate 1 white-background hero + 3 alpine scenes + 2 desert scenes, estimated 12 credits, each anchored to user upload."

Three choices for the user: approve, cancel, or refine in natural language. No credits are spent until approval.

  • •Zero credits consumed before approval — no generation calls during planning
  • •Iterative revision — refine the plan as many times as needed
  • •Approved plans are archived for full credit traceability after the fact
05

Intelligent Result Cache

Defends against model-induced redundancy · Tool Result Cache

Mid-tier LLMs share a familiar flaw: they invoke the same tool twice, sometimes three times, burning tokens each round.

We reuse results within a single run — repeated read-only operations never burn compute twice; the prior result is reused instantly.

  • •Repeated read-only operations auto-reuse their result within a run
  • •Duplicate calls are recognized precisely, avoiding wasted re-execution
  • •Complex tasks see a significant reduction in token spend
06

Persistent Long-Session Memory

Two-tier memory architecture · Context Memory

A single apparel shoot can produce hundreds of images across dozens of conversation turns. Most frameworks either blow the context window or start hallucinating.

Two memory layers: short-term auto-summarization, and long-term persistence of the essentials.

  • •Short-term — as a conversation grows, early turns are auto-summarized to free up context
  • •Long-term — key information is remembered persistently and carried across sessions
  • •Every image carries its provenance (who uploaded it, which turn, what was asked) — no mix-ups
  • •At turn 80, the agent still remembers what was uploaded at turn 3
07

Crash-Resilient Task System

Incremental IDs, positional indexing, local persistence · Task Persistence

In enterprise environments, crashes, power loss, and forced restarts are routine. Most AI tools lose the in-flight task entirely.

Our task system uses a three-layer design that recovers from crashes with zero data loss.

  • •Task IDs are incremental strings (1 / 2 / 3) rather than UUIDs — lighter cognitive load for the model, more reliable dispatch
  • •Subtasks use positional indices (0 / 1 / 2) so the model never has to memorize long strings
  • •Up to 100 tasks are persisted locally for inspection and recovery at any time
  • •Auto-repair on restart — orphaned "in-progress" tasks are demoted to paused, eliminating ghost jobs
08

Network Fault Tolerance

API jitter doesn't break the render · Network Resilience

Anyone shipping against third-party LLM APIs knows the drill — occasional timeouts, 500s, and rate limits. The agent has to absorb all of it.

  • •120s API timeout, generous enough for extended-thinking models
  • •3 retries with exponential backoff (500ms → 1s → 2s)
  • •4xx client errors fail fast; 5xx, 429, and timeouts auto-retry
  • •Malformed responses are treated as failures and retried rather than passed through as empty
  • •Empty responses surface explicit errors (likely safety filtering or max_tokens exhausted by thinking) instead of silent exits
09

Dynamic System Prompt

Always grounded in the live canvas · Dynamic Prompt

Most agents bake the system prompt once at startup — they never see that the canvas changed, the model swapped, or new references were uploaded.

Our agent rebuilds the system prompt every iteration, injecting live canvas state, reference pool, current selection, and model capabilities.

  • •Real-time awareness of canvas contents, aspect ratio, and resolution
  • •Reference pool overview is indexed per-slot so the agent never mixes them up
  • •Available models and current-model capabilities (multi-reference, inpainting, max-N) are injected live
  • •Switch models and the agent immediately knows the new capability surface
10

Full Event Transparency

No black box at any step · Event Stream

What the agent is thinking, which tool it invoked, what came back, why it's requesting confirmation — every signal streams to the UI in real time.

Clients see every decision the agent makes — a completely different trust profile from the "spinner-until-result" black-box tools.

  • •Event types — thinking, tool_call, tool_result, message, permission_request, error
  • •Tool-call arguments are exposed live, letting clients reverse-engineer the agent's reasoning
  • •Errors are humanized into actionable advice — switch model, simplify request, and so on