Skip to content

Image describe ​

  • Category: Generate

What it is ​

Sends one to eight local still frames to the official still-frame describe service and writes a structured result (JSON Document v2), a summary, and an optional Per-frame results array. This node only produces the description; what it is used for (score / SFX / search and so on) is up to the downstream node. It has no built-in sound search; to search for sounds, connect the summary or a search phrase to "Find sounds".

Describe modes ​

ModeBehaviorBilling
Combined (default)One vision-language call for all images, one combined description1 invoke (more images raise the tokens of that one call; not ×n)
Per frameOne caption per image (sent one by one; the gateway has no batch result)N invokes (N ≤ 8)

Taking frames upstream is free on your computer.

Inputs ​

NameDescription
Image inputLocal still frames or frames grabbed from the timeline, up to 8

This node has no frame metadata input port. When the upstream node is "Get image", the run silently merges its frame positions into sourceFrames of the Document and of the per-frame output, so no extra wiring is needed.

Outputs ​

NameDescription
ResultDocument: schemaVersion: 2, result, sourceFrames (in per-frame mode each frame can carry a caption), optional queryForSearch
SummaryCombined: result.summary. Per frame: the captions of all frames joined by line breaks
Per-frame resultsAn array in the same shape as the frame metadata of "Get image", which can carry a caption per frame. Connect it to the frame metadata of "Create timeline markers" to place a point per frame

Settings ​

NameDescription
Describe modeCombined / Per frame (segmented bar; stored in configValues.describeMode)

The purpose required by the gateway contract is fixed to generic by the client and cannot be chosen.

Typical wiring ​

ScenarioHow to connect
Frame → describe"Get image" image output → this node's image input (positions go into the result silently)
Combined → markersThis node's result → "Create timeline markers" frame metadata (all points share the combined summary)
Per frame → markersThis node's Per-frame results → "Create timeline markers" frame metadata (one caption per point)
Frame → positions only"Get image" frame metadata → "Create timeline markers" frame metadata (no need to describe first)
Text onlySummary → "Find sounds" / "Text generation"

Typical downstream ​

Text generation, Find sounds, Create timeline markers

Notes ​

  • At most 8 images per run; extra images are cut off.
  • To use the frame positions on their own, wire the frame metadata output of "Get image", or this node's "Per-frame results".
  • On failure, the existing short messages for sign-in / account benefits / caps are shown. In per-frame mode, if any frame fails the whole run fails.
  • The result card still shows the summary; expand the preview to see the full Document (including sourceFrames).

Agents / MCP ​

For agents or developers (optional)
  • Node type: still-describe
  • Input port: image-in only; output ports: describe-io, summary, frames-describe-io
  • Setting: describeMode = combined | per_frame (default combined); the gateway purpose is fixed to generic in the run layer
  • Document: { schemaVersion: 2, result, sourceFrames, queryForSearch? }; a frame can carry a caption
  • frames-describe-io: BlueprintHostImageSourceFrame[] (same shape as get-image frames-meta)
  • Silent pass-through: upstream host-timeline-get-image on image-in → timeline is merged
  • Billing: combined 1 invoke; per frame N invokes
  • Typical upstream: host-timeline-get-image; downstream: generate-text, audio-search, marker-create