| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
| Name | Name | Last commit date | ||
|---|---|---|---|---|
Your GPU, as something an agent can drive.
diffusers-workflow wraps the Hugging Face Diffusers library in an engine that runs image, video and audio generation as jobs, and puts two front ends on it: an MCP server, so Claude Code (or any MCP client) can author, run and inspect generations; and a web UI for doing the same by hand. A CLI sits underneath for when you want neither.
Python 3.10-3.14 | CUDA (NVIDIA) | MPS (Apple Silicon) | CPU
1. Install. The script picks the right torch build for your platform, creates a virtual environment and installs everything, MCP server included.
# Linux / macOS
bash ./install.sh
source ./activate
# Windows
.\install.ps1
.\venv\scripts\activatepython -m dw.test runs a small built-in workflow end to end - it downloads SD 1.5 (a few GB) and generates one image on whichever accelerator was found.
2. Start the engine. Leave it running; everything else talks to it.
dw-serve
# diffusers-workflow server on http://127.0.0.1:8765That address is the web UI. Open it and run templates/text-to-image — a small, ungated model, so the first generation needs no Hugging Face login and downloads only a few GB.
3. Connect Claude Code. Register the MCP server with the absolute path to dw-mcp in the venv you just made (the relative path is the one setup detail that reliably goes wrong):
claude mcp add dw -- "$(pwd)/venv/bin/dw-mcp"Then, optionally, the dw plugin — one skill per model family that knows which workflow fits a request and the rules that bite:
/plugin marketplace add dkackman/diffusers-workflow /plugin install dw@diffusers-workflow
Most of the shipped workflows (Flux, LTX-2, MiniMax...) use gated models. Request access on the model's Hugging Face page, then huggingface-cli login once; without it the run fails partway through with a 401/403 from the Hub.
GPU on another machine? Start the engine there with --mcp and connect over HTTP — nothing to install on the laptop:
# on the GPU box dw-serve --host 0.0.0.0 --token "$DW_API_TOKEN" --mcp --workspace ~/studio # on your laptop claude mcp add --transport http dw http://gpu-box:8765/mcp \ --header "Authorization: Bearer $DW_API_TOKEN"The server's own Server page composes that line for the address you pick. End to end: Remote GPU server.
Then just ask. The agent has 62 tools covering the whole surface — the workflow catalog, the real diffusers pipeline signatures, the job queue, the gallery, the model cache:
Generation is the long pass, and the agent stays with it — queuing each shot, waiting it out, and reporting what came back:
What a session looks like:
Everything that costs real GPU time or real disk (run_workflow, rerun_job, enhance_prompt, download_model, delete_model, update_diffusers, delete_workspace) refuses until it is explicitly acknowledged, so an agent cannot quietly burn an hour of GPU or delete 40GB of weights.
One server holds several workspaces — each with its own workflows, assets and outputs — so two agents, or an agent and you in the browser, share the GPU without saving over each other. An agent calls use_workspace once and the rest of the session lands there.
Feedback from a session. At the end of a working session, ask the agent what got in its way: bugs, gaps, misleading skill text, tools it reached for and couldn't find. Have it file each one as an issue on dkackman/diffusers-workflow with the field-report label, for example:
File each bug or gap you hit as an issue on dkackman/diffusers-workflow with the label field-report.
The label marks a report as coming from real use, not from the automated test loop. The agent loop (see Agent Loop) picks the report up like any other issue when you filed it yourself; a report filed under any other GitHub login is parked for the maintainer to review first (relabelled owner:don + status:needs-approval), since the loop must not act unattended on third-party text in a public repo, and is only handed to the loop, or not, after that review. The label is also what feature planning reads as evidence of demand.
The complete tool reference, client configuration for other MCP hosts, and the troubleshooting table: MCP Server. Workspaces in depth: Workspaces.
Everything the engine does, in a browser, backed by the same persistent GPU worker — models stay loaded between runs.
An editor built from the real pipeline signatures. Forms and argument autocomplete are generated by introspecting diffusers itself, so every knob a pipeline exposes is there with its documentation. Validation catches schema errors and argument typos before any model loads.
A gallery where every image is a recipe. Outputs carry their full workflow and seed; open as workflow drops any image back into the editor, ready to reproduce or riff on. Keep as asset promotes a generated file into the asset library for later workflows to build on.
A prompt library stores a prompt once and lets any workflow reference it, with an Enhance with AI panel that expands an idea into a full prompt using a local language model. A model manager inventories the Hugging Face hub cache — sizes, last use, free space — and downloads or deletes models with live progress.
Jobs queue, stream progress live per denoising step, cancel cooperatively and persist to a searchable history. See Server & Web UI for the pages and the HTTP API.
dw.run is a thin client of dw.serve: it queues a job over HTTP, prints its progress, and reports the run directory the server wrote to - start the server first.
python -m dw.serve
python -m dw.run workflows/templates/text-to-image.json
python -m dw.run workflows/templates/text-to-image.json prompt="a cat" num_images_per_prompt=4
python -m dw.validate workflows/models/flux-dev.jsondw.run takes WORKFLOW [name=value ...] [--server URL] [--workspace NAME] [--token TOKEN] - --server defaults to DW_MCP_URL, else http://127.0.0.1:8765; --token to DW_API_TOKEN. Where the output lands, which workspace's prompts:/asset: references resolve, and the output layout are all dw.serve flags now (--workspace, --output-dir, --prompt-dir, --asset-dir, --output-layout on the server) - dw.run only says which server and which of its workspaces to run in.
Every front end reads and writes the same thing: a JSON document of named steps, each a diffusers pipeline or a utility task, whose arguments reference variables, earlier steps' outputs, stored prompts and assets rather than hard-coded values. That is what makes text-to-image chain into image-to-video, and what makes a generated image reopen as the exact recipe that produced it. workflows/ is a corpus of runnable examples across model families; the Workflow Guide is the reference when you do want to write one.
Because a workflow reaches any diffusers pipeline or quantization backend by dynamic import, loading one can execute arbitrary Python. Treat a workflow file from someone else the way you'd treat a .py script — see Trust model.
Under the hood the engine also handles: quantization (BitsAndBytes, TorchAO, GGUF, SDNQ, optimum-quanto); inference acceleration (FirstBlockCache, FasterCache, MagCache, TaylorSeerCache); LoRA and IP-Adapter; A1111-style prompt weighting; long-video chaining with audio-driven length; step-output caching, so re-running a fixed-seed workflow finishes instantly; and utility tasks for upscaling, face restoration, segmentation, captioning, frame interpolation and more.
| Back | FazBrowse Home | New Git URL |