Image describe
- Category: Generate
What it is
Sends one to eight local still frames to the official still-frame describe service and writes a structured result (JSON Document v2), a summary, and an optional Per-frame results array. This node only produces the description; what it is used for (score / SFX / search and so on) is up to the downstream node. It has no built-in sound search; to search for sounds, connect the summary or a search phrase to "Find sounds".
Describe modes
| Mode | Behavior | Billing |
|---|---|---|
| Combined (default) | One vision-language call for all images, one combined description | 1 invoke (more images raise the tokens of that one call; not ×n) |
| Per frame | One caption per image (sent one by one; the gateway has no batch result) | N invokes (N ≤ 8) |
Taking frames upstream is free on your computer.
Inputs
| Name | Description |
|---|---|
| Image input | Local still frames or frames grabbed from the timeline, up to 8 |
This node has no frame metadata input port. When the upstream node is "Get image", the run silently merges its frame positions into sourceFrames of the Document and of the per-frame output, so no extra wiring is needed.
Outputs
| Name | Description |
|---|---|
| Result | Document: schemaVersion: 2, result, sourceFrames (in per-frame mode each frame can carry a caption), optional queryForSearch |
| Summary | Combined: result.summary. Per frame: the captions of all frames joined by line breaks |
| Per-frame results | An array in the same shape as the frame metadata of "Get image", which can carry a caption per frame. Connect it to the frame metadata of "Create timeline markers" to place a point per frame |
Settings
| Name | Description |
|---|---|
| Describe mode | Combined / Per frame (segmented bar; stored in configValues.describeMode) |
The purpose required by the gateway contract is fixed to generic by the client and cannot be chosen.
Typical wiring
| Scenario | How to connect |
|---|---|
| Frame → describe | "Get image" image output → this node's image input (positions go into the result silently) |
| Combined → markers | This node's result → "Create timeline markers" frame metadata (all points share the combined summary) |
| Per frame → markers | This node's Per-frame results → "Create timeline markers" frame metadata (one caption per point) |
| Frame → positions only | "Get image" frame metadata → "Create timeline markers" frame metadata (no need to describe first) |
| Text only | Summary → "Find sounds" / "Text generation" |
Typical downstream
Text generation, Find sounds, Create timeline markers
Notes
- At most 8 images per run; extra images are cut off.
- To use the frame positions on their own, wire the frame metadata output of "Get image", or this node's "Per-frame results".
- On failure, the existing short messages for sign-in / account benefits / caps are shown. In per-frame mode, if any frame fails the whole run fails.
- The result card still shows the summary; expand the preview to see the full Document (including
sourceFrames).
Agents / MCP
For agents or developers (optional)
- Node type:
still-describe - Input port:
image-inonly; output ports:describe-io,summary,frames-describe-io - Setting:
describeMode=combined|per_frame(default combined); the gatewaypurposeis fixed togenericin the run layer - Document:
{ schemaVersion: 2, result, sourceFrames, queryForSearch? }; a frame can carry acaption frames-describe-io:BlueprintHostImageSourceFrame[](same shape as get-imageframes-meta)- Silent pass-through: upstream
host-timeline-get-imageon image-in → timeline is merged - Billing: combined 1 invoke; per frame N invokes
- Typical upstream:
host-timeline-get-image; downstream:generate-text,audio-search,marker-create