Product craft · Field note

A reviewable multimodal production pipeline

How I organise generated 2D, 3D, audio, voice, and video as a traceable production process rather than a collection of one-off prompts.

  • ComfyUI
  • Multimodal AI
  • Asset pipelines
  • Human review

I treat generated media as an intermediate artefact. It still has to meet the product brief, technical constraints, continuity requirements, and rights checks before it belongs in a release.

Turn experiments into stages

My experiments cover 2D, 3D, and isometric image work through ComfyUI and custom diffusion training, alongside audio, voice, and video. Once an approach is useful, I turn it into a small pipeline with a defined input, an inspectable output, and a named owner for the next decision.

That removes a common source of hidden work: a promising asset that nobody can reproduce, revise, or license with confidence.

Keep provenance beside the asset

Each stage records the source material, model choice, transformations, and review state. The record does not need to be elaborate; it needs to answer practical questions later:

  • Which inputs and rights were used?
  • Can the result be regenerated or edited?
  • What product and technical checks has it passed?
  • Who approved it for release?

Continuity, rights, technical fit, and human review remain release gates. A polished image is not automatically a shippable asset.

Route for fit, not novelty

The pipeline routes work according to quality, latency, and cost rather than sending every request to the largest model. Some stages need a specialist model; others are more reliable as deterministic tools or ordinary manual work.

Because the boundaries are explicit, I can replace a model without rebuilding the production loop. That matters in a landscape where capabilities, prices, and licence terms change quickly.