| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
| Name | Name | Last commit date | ||
|---|---|---|---|---|
parent directory.. | ||||
Example/demo server that keeps a single model in memory while safely running parallel inference requests by creating per-request lightweight views and cloning only small, stateful components (schedulers, RNG state, small mutable attrs). Works with StableDiffusion3 pipelines. We recommend running 10 to 50 inferences in parallel for optimal performance, averaging between 25 and 30 seconds to 1 minute and 1 minute and 30 seconds. (This is only recommended if you have a GPU with 35GB of VRAM or more; otherwise, keep it to one or two inferences in parallel to avoid decoding or saving errors due to memory shortages.)
All the components needed to create the inference server are in the current directory:
server-async/ ├── utils/ ├─────── __init__.py ├─────── scheduler.py # BaseAsyncScheduler wrapper and async_retrieve_timesteps for secure inferences ├─────── requestscopedpipeline.py # RequestScoped Pipeline for inference with a single in-memory model ├─────── utils.py # Image/video saving utilities and service configuration ├── Pipelines.py # pipeline loader classes (SD3) ├── serverasync.py # FastAPI app with lifespan management and async inference endpoints ├── test.py # Client test script for inference requests ├── requirements.txt # Dependencies └── README.md # This documentation
Core problem: a naive server that calls pipe.__call__ concurrently can hit race conditions (e.g., scheduler.set_timesteps mutates shared state) or explode memory by deep-copying the whole pipeline per-request.
diffusers-async / this example addresses that by:
Single model instance is loaded into memory (GPU/MPS) when the server starts.
On each HTTP inference request:
The server uses RequestScopedPipeline.generate(...) which:
Result: inference completes, images are moved to CPU & saved (if requested), internal buffers freed (GC + torch.cuda.empty_cache()).
Multiple requests can run in parallel while sharing heavy weights and isolating mutable state.
Recommended: create a virtualenv / conda environment.
pip install diffusers
pip install -r requirements.txtUsing the serverasync.py file that already has everything you need:
python serverasync.pyThe server will start on http://localhost:8500 by default with the following features:
Use the included test script:
python test.pyOr send a manual request:
POST /api/diffusers/inference with JSON body:
{
"prompt": "A futuristic cityscape, vibrant colors",
"num_inference_steps": 30,
"num_images_per_prompt": 1
}Response example:
{
"response": ["http://localhost:8500/images/img123.png"]
}RequestScopedPipeline(
pipeline, # Base pipeline to wrap
mutable_attrs=None, # Custom list of attributes to clone
auto_detect_mutables=True, # Enable automatic detection of mutable attributes
tensor_numel_threshold=1_000_000, # Tensor size threshold for cloning
tokenizer_lock=None, # Custom threading lock for tokenizers
wrap_scheduler=True # Auto-wrap scheduler in BaseAsyncScheduler
)The server configuration can be modified in serverasync.py through the ServerConfigModels dataclass:
@dataclass
class ServerConfigModels:
model: str = 'stabilityai/stable-diffusion-3.5-medium'
type_models: str = 't2im'
host: str = '0.0.0.0'
port: int = 8500Already borrowed — previously a Rust tokenizer concurrency error. ✅ This is now fixed: RequestScopedPipeline automatically detects and wraps tokenizers with thread locks, so race conditions no longer happen.
can't set attribute 'components' — pipeline exposes read-only components. ✅ The RequestScopedPipeline now detects read-only properties and skips setting them automatically.
Scheduler issues:
Memory issues with large tensors: ✅ The system now has configurable tensor_numel_threshold to prevent cloning of large tensors while still cloning small mutable ones.
Automatic tokenizer detection: ✅ The system automatically identifies tokenizer components by checking for tokenizer methods, class names, and attributes, then applies thread-safe wrappers.
| Back | FazBrowse Home | New Git URL |