From 46673b422450db615324e85a64b35b4b3a960f08 Mon Sep 17 00:00:00 2001 From: Mick Date: Wed, 26 Nov 2025 00:02:14 +0800 Subject: [PATCH] [diffusion] doc: add doc for LoRA usage (#13931) --- python/sglang/multimodal_gen/docs/cli.md | 94 +----- .../sglang/multimodal_gen/docs/openai_api.md | 281 ++++++++++++++++++ 2 files changed, 283 insertions(+), 92 deletions(-) create mode 100644 python/sglang/multimodal_gen/docs/openai_api.md diff --git a/python/sglang/multimodal_gen/docs/cli.md b/python/sglang/multimodal_gen/docs/cli.md index e9471c593..4b0f29c72 100644 --- a/python/sglang/multimodal_gen/docs/cli.md +++ b/python/sglang/multimodal_gen/docs/cli.md @@ -127,7 +127,7 @@ sglang generate --help ## Serve -Launch the SGLang diffusion HTTP server and interact with it using the OpenAI SDK and curl. The server implements an OpenAI-compatible subset for Videos under the `/v1/videos` namespace. +Launch the SGLang diffusion HTTP server and interact with it using the OpenAI SDK and curl. ### Start the server @@ -149,98 +149,8 @@ sglang serve "${SERVER_ARGS[@]}" - **--model-path**: Which model to load. The example uses `Wan-AI/Wan2.1-T2V-1.3B-Diffusers`. - **--port**: HTTP port to listen on (the default here is `30010`). -Wait until the port is listening. In CI, the tests probe `127.0.0.1:30010` before sending requests. +For detailed API usage, including Image, Video Generation and LoRA management, please refer to the [OpenAI API Documentation](openai_api.md). -### OpenAI Python SDK usage - -Initialize the client with a dummy API key and point `base_url` to your local server: - -```python -from openai import OpenAI - -client = OpenAI(api_key="sk-proj-1234567890", base_url="http://localhost:30010/v1") -``` - -- **Create a video** - -```python -video = client.videos.create(prompt="A calico cat playing a piano on stage", size="1280x720") -print(video.id, video.status) -``` - -Response example fields include `id`, `status` (e.g., `queued` → `completed`), `size`, and `seconds`. - -- **List videos** - -```python -videos = client.videos.list() -for item in videos.data: - print(item.id, item.status) -``` - -- **Poll for completion and download content** - -```python -import time - -video = client.videos.create(prompt="A calico cat playing a piano on stage", size="1280x720") -video_id = video.id - -# Simple polling loop -while True: - page = client.videos.list() - item = next((v for v in page.data if v.id == video_id), None) - if item and item.status == "completed": - break - time.sleep(5) - -# Download binary content (MP4) -resp = client.videos.download_content(video_id=video_id) -content = resp.read() # bytes -with open("output.mp4", "wb") as f: - f.write(content) -``` - -### curl examples - -- **Create a video** - -```bash -curl -sS -X POST "http://localhost:30010/v1/videos" \ - -H "Content-Type: application/json" \ - -H "Authorization: Bearer sk-proj-1234567890" \ - -d '{ - "prompt": "A calico cat playing a piano on stage", - "size": "1280x720" - }' -``` - -- **List videos** - -```bash -curl -sS -X GET "http://localhost:30010/v1/videos" \ - -H "Authorization: Bearer sk-proj-1234567890" -``` - -- **Download video content** - -```bash -curl -sS -L "http://localhost:30010/v1/videos//content" \ - -H "Authorization: Bearer sk-proj-1234567890" \ - -o output.mp4 -``` - -### API surface implemented here - -The server exposes these endpoints (OpenAPI tag `videos`): - -- `POST /v1/videos` — Create a generation job and return a queued `video` object. -- `GET /v1/videos` — List jobs. -- `GET /v1/videos/{video_id}/content` — Download binary content when ready (e.g., MP4). - -### Reference - -- OpenAI Videos API reference: `https://platform.openai.com/docs/api-reference/videos` ## Generate diff --git a/python/sglang/multimodal_gen/docs/openai_api.md b/python/sglang/multimodal_gen/docs/openai_api.md new file mode 100644 index 000000000..568104f13 --- /dev/null +++ b/python/sglang/multimodal_gen/docs/openai_api.md @@ -0,0 +1,281 @@ +# SGLang Diffusion OpenAI API + +The SGLang diffusion HTTP server implements an OpenAI-compatible API for image and video generation, as well as LoRA adapter management. + +## Serve + +Launch the server using the `sglang serve` command. + +### Start the server + +```bash +SERVER_ARGS=( + --model-path Wan-AI/Wan2.1-T2V-1.3B-Diffusers + --text-encoder-cpu-offload + --pin-cpu-memory + --num-gpus 4 + --ulysses-degree=2 + --ring-degree=2 + --port 30010 +) + +sglang serve "${SERVER_ARGS[@]}" +``` + +- **--model-path**: Path to the model or model ID. +- **--port**: HTTP port to listen on (default: `30000`). + +--- + +## Endpoints + +### Image Generation + +The server implements an OpenAI-compatible Images API under the `/v1/images` namespace. + +#### Create an image + +**Endpoint:** `POST /v1/images/generations` + +**Python Example (b64_json response):** + +```python +import base64 +from openai import OpenAI + +client = OpenAI(api_key="sk-proj-1234567890", base_url="http://localhost:30010/v1") + +img = client.images.generate( + prompt="A calico cat playing a piano on stage", + size="1024x1024", + n=1, + response_format="b64_json", +) + +image_bytes = base64.b64decode(img.data[0].b64_json) +with open("output.png", "wb") as f: + f.write(image_bytes) +``` + +**Curl Example:** + +```bash +curl -sS -X POST "http://localhost:30010/v1/images/generations" \ + -H "Content-Type: application/json" \ + -H "Authorization: Bearer sk-proj-1234567890" \ + -d '{ + "prompt": "A calico cat playing a piano on stage", + "size": "1024x1024", + "n": 1, + "response_format": "b64_json" + }' +``` + +> **Note** +> The `response_format=url` option is not supported for `POST /v1/images/generations` and will return a `400` error. + +#### Edit an image + +**Endpoint:** `POST /v1/images/edits` + +This endpoint accepts a multipart form upload with an input image and a text prompt. The server can return either a base64-encoded image or a URL to download the image. + +**Curl Example (b64_json response):** + +```bash +curl -sS -X POST "http://localhost:30010/v1/images/edits" \ + -H "Authorization: Bearer sk-proj-1234567890" \ + -F "image=@input.png" \ + -F "prompt=A calico cat playing a piano on stage" \ + -F "size=1024x1024" \ + -F "response_format=b64_json" +``` + +**Curl Example (URL response):** + +```bash +curl -sS -X POST "http://localhost:30010/v1/images/edits" \ + -H "Authorization: Bearer sk-proj-1234567890" \ + -F "image=@input.png" \ + -F "prompt=A calico cat playing a piano on stage" \ + -F "size=1024x1024" \ + -F "response_format=url" +``` + +#### Download image content + +When `response_format=url` is used with `POST /v1/images/edits`, the API returns a relative URL like `/v1/images//content`. + +**Endpoint:** `GET /v1/images/{image_id}/content` + +**Curl Example:** + +```bash +curl -sS -L "http://localhost:30010/v1/images//content" \ + -H "Authorization: Bearer sk-proj-1234567890" \ + -o output.png +``` + +### Video Generation + +The server implements a subset of the OpenAI Videos API under the `/v1/videos` namespace. + +#### Create a video + +**Endpoint:** `POST /v1/videos` + +**Python Example:** + +```python +from openai import OpenAI + +client = OpenAI(api_key="sk-proj-1234567890", base_url="http://localhost:30010/v1") + +video = client.videos.create( + prompt="A calico cat playing a piano on stage", + size="1280x720" +) +print(f"Video ID: {video.id}, Status: {video.status}") +``` + +**Curl Example:** + +```bash +curl -sS -X POST "http://localhost:30010/v1/videos" \ + -H "Content-Type: application/json" \ + -H "Authorization: Bearer sk-proj-1234567890" \ + -d '{ + "prompt": "A calico cat playing a piano on stage", + "size": "1280x720" + }' +``` + +#### List videos + +**Endpoint:** `GET /v1/videos` + +**Python Example:** + +```python +videos = client.videos.list() +for item in videos.data: + print(item.id, item.status) +``` + +**Curl Example:** + +```bash +curl -sS -X GET "http://localhost:30010/v1/videos" \ + -H "Authorization: Bearer sk-proj-1234567890" +``` + +#### Download video content + +**Endpoint:** `GET /v1/videos/{video_id}/content` + +**Python Example:** + +```python +import time + +# Poll for completion +while True: + page = client.videos.list() + item = next((v for v in page.data if v.id == video_id), None) + if item and item.status == "completed": + break + time.sleep(5) + +# Download content +resp = client.videos.download_content(video_id=video_id) +with open("output.mp4", "wb") as f: + f.write(resp.read()) +``` + +**Curl Example:** + +```bash +curl -sS -L "http://localhost:30010/v1/videos//content" \ + -H "Authorization: Bearer sk-proj-1234567890" \ + -o output.mp4 +``` + +--- + +### LoRA Management + +The server supports dynamic loading, merging, and unmerging of LoRA adapters. + +**Important Notes:** +- Mutual Exclusion: Only one LoRA can be *merged* (active) at a time +- Switching: To switch LoRAs, you must first `unmerge` the current one, then `set` the new one +- Caching: The server caches loaded LoRA weights in memory. Switching back to a previously loaded LoRA (same path) has little cost + +#### Set LoRA Adapter + +Loads a LoRA adapter and merges its weights into the model. + +**Endpoint:** `POST /v1/set_lora` + +**Parameters:** +- `lora_nickname` (string, required): A unique identifier for this LoRA +- `lora_path` (string, optional): Path to the `.safetensors` file or Hugging Face repo ID. Required for the first load; optional if re-activating a cached nickname + +**Curl Example:** + +```bash +curl -X POST http://localhost:30010/v1/set_lora \ + -H "Content-Type: application/json" \ + -d '{ + "lora_nickname": "lora_name", + "lora_path": "/path/to/lora.safetensors" + }' +``` + + +#### Merge LoRA Weights + +Manually merges the currently set LoRA weights into the base model. + +> [!NOTE] +> `set_lora` automatically performs a merge, so this is typically only needed if you have manually unmerged but want to re-apply the same LoRA without calling `set_lora` again.* + +**Endpoint:** `POST /v1/merge_lora_weights` + +**Curl Example:** + +```bash +curl -X POST http://localhost:30010/v1/merge_lora_weights \ + -H "Content-Type: application/json" +``` + + +#### Unmerge LoRA Weights + +Unmerges the currently active LoRA weights from the base model, restoring it to its original state. This **must** be called before setting a different LoRA. + +**Endpoint:** `POST /v1/unmerge_lora_weights` + +**Curl Example:** + +```bash +curl -X POST http://localhost:30010/v1/unmerge_lora_weights \ + -H "Content-Type: application/json" +``` + +### Example: Switching LoRAs + +1. Set LoRA A: + ```bash + curl -X POST http://localhost:30010/v1/set_lora -d '{"lora_nickname": "lora_a", "lora_path": "path/to/A"}' + ``` +2. Generate with LoRA A... +3. Unmerge LoRA A: + ```bash + curl -X POST http://localhost:30010/v1/unmerge_lora_weights + ``` +4. Set LoRA B: + ```bash + curl -X POST http://localhost:30010/v1/set_lora -d '{"lora_nickname": "lora_b", "lora_path": "path/to/B"}' + ``` +5. Generate with LoRA B...