Skip to content
← All work
2023–2026In progressRuns locally

Random Acts of Creativity with AI

Three years at the frontier of generative image, video and audio: from hand-built ComfyUI experiments, through a platform that licensed my work, to a local pipeline where project agents render art, film, voice and music alongside the rest of the build.

Director and toolsmith: workflows, MCP servers, pipelines, art direction

At the frontier since
Nov 2023
Horiar era
202 commits · 224 clips
ComfyUI workflows as tools
17
MCP servers written
~20

01 / problem

The problem

I've been exploring generative image and video since late 2023, and the frontier keeps moving. Every few months a new model makes last season's workflow obsolete.

The interesting problem isn't any single model. It's building a pipeline that gets better with each release instead of starting over, and that agents can eventually drive for themselves.

02 / build

What I built

Six stages, each one carrying the best idea of the last:

  • Frontier, by hand (2023–24). ComfyUI animation, AI dream sequences, image-to-3D, local models driving Stable Diffusion, and published ComfyUI workflows, all shared as I went on YouTube (below). The first Sephira songs came from this era: Suno for the music, ElevenLabs for her voice, early SDXL for the visuals.
  • Kinetic Canvas, with Horiar (Dec 2024 to Mar 2025). Horiar, a team building a generative platform, saw the work as artistic research and licensed their product to me to mature it. I built Kinetic Canvas on top: an LLM that structures every prompt (493 logged), video extended by pulling each clip's last frame into the next, and an ffmpeg timeline to cut it together. 202 commits and 224 clips, about three hours of footage.
  • A local stack (2025). Imagen 4 and Veo 3 through my own command-line tool to learn what top quality looked like, then the first local ComfyUI video, and the first MCP server that let an agent call Flux and Qwen.
  • A workflow catalog for agents (2026). On a new GPU, ComfyUI became 17 typed tools in an MCP server I wrote: Qwen image, edit and ControlNet, WAN 2.2 image-to-video and first-to-last-frame video, upscaling and image-to-3D. Alongside it, MCP servers for ACE-Step music and Chatterbox voice. Any agent can now render image, video, 3D, music and voice. Sephira was re-produced locally with ACE-Step from my own lyrics (listen below).
  • vidBridge (summer 2026). The primitives became a directed film tool: a one-line spark becomes an LLM storyboard; Qwen paints each shot's first and last frames; WAN 2.2 bridges them; a vision model reads the last frame to write the next shot. It fixed the drift of the old last-frame chain and keeps characters consistent. Eight films so far, including the one above.
  • Assets as code (now). Project agents generate inside product repos, next to ffmpeg, sharp, opentype and Playwright. Lull's eleven paintings and its Halloween sound are npm run steps. Waywyrd's teaser is rendered frame by frame with its own sound design. Biome's broadcast runs on 118 generated persona clips. Ninth Candle paints its scenes at runtime.

03 / signal

What it shows

The models changed every few months; the pipeline compounded. Ideas carried forward from stage to stage: an LLM structuring the prompt, chaining shots by their frames, a director's approval before anything expensive renders. Eventually generation stopped being a separate activity and became an ordinary build step, and then a feature products use at runtime.

It also shows the director's side: knowing what to make, judging what's good, and building the tools so the next idea costs less than the last one.

Gallery 1 / 6

“Lines of Consciousness”, made in vidBridge: an LLM wrote the storyboard, Qwen painted each shot’s first and last frames, and WAN 2.2 bridged them.

Listen

Sephira: “Not Your Basic”

2024 originalSuno

2026 remasterACE-Step, local, from my lyrics

Sephira: “(Like a) Natural Language”

2024 originalSuno

2026 remixACE-Step, local, from my lyrics

Waywyrd: ritual audio

ChantACE-Step, generated to brief

Lyre bedACE-Step, generated to brief

Paired clips are 45-second excerpts taken from the same point in each version.

On video

Alien Crab: Animation in ComfyUI2023

Early ComfyUI animation, from the first stage.

The SubSymph Workflow (V1)2024

A published ComfyUI animation workflow, distilled from many hours of testing.

(Like a) Natural Language2024

The original Sephira release: Suno music, ElevenLabs voice, early SDXL visuals.

Kinetic Canvas: a custom Horiar text-to-video workflow2025

The Horiar stage: structured prompts and chained clips.