CodeAFPictures, video, voice and music

Pictures, video, voice and music

Ask for a picture, a clip, a voiceover or a jingle in plain words, and get the file.

Media is one more thing the agent can make, with the same words, the same budget and the same files as code.
  • draw a 16:9 banner for the release
  • a 5 second clip of the product in use, 1080p
  • read the changelog aloud, and a 10 second jingle under it

CodeAF makes four kinds of media: pictures, video, spoken audio and music. Each kind is a tool the agent calls in the middle of its work, each runs on its own model, and each saves a real file in your project and names the path. The agent writes the prompt from your words, and it can look at, listen to or frame-grab what it made and try again before it calls the job done.

Try it in CodeAF

Where

In a conversation in your repo.

Type
Draw a 16:9 banner for the next release, flat two colour print on warm paper, and save it as media/banner.png
You see

A step that says it is drawing the banner, then an answer naming the picture's full path, media/banner.png in your repo.

Four kinds, four models

KindDefault modelOptions you can name
PicturesSeedream 5.0 Proaspect ratio, size, pictures to edit or combine
VideoSeedance 2.5length, aspect ratio, resolution, seed, first and last frames
SpeechFish Audio S2.1 Provoice
MusicLyria 3 Clipmood, instruments and length, in the brief

You ask in words: "draw…", "make a 5 second clip of…", "read this aloud…", "compose a jingle…". From a script, codeaf image "prompt" --out file.png makes one picture and exits; it is the one media command on the command line.

In our run a 2048 pixel picture cost 9 cents, a 5 second 480p clip 21 cents, a jingle 6 cents for the whole turn and a short voiceover under one cent. Pictures and speech arrive in the same turn; video and music render in the background and arrive as a note a minute or two later. /files lists everything made for you, newest first.

Choose the model

Every kind has its own row in /settings, on the Providers tab: drawing, filming, speaking and composing, each automatic until you pick one. Press Enter on a row for the list of models that make that kind, with prices and scores, type to filter, and press Enter again; the choice is saved to your profile. The lists are long: more than 50 image models (Flux, Gemini, GPT Image, Recraft, Seedream, Krea), close to 30 video models (Veo, Kling, Sora, Runway, Seedance, Wan, Hailuo) and about 20 voices (Fish Audio, Gemini TTS, MiniMax, Deepgram, Kokoro). The composing list comes up empty in the current build, so music stays on Lyria unless you name a model or set the variable below. Scripts and servers can set the same thing with CODEAF_IMAGE_MODEL, CODEAF_VIDEO_MODEL, CODEAF_SPEECH_MODEL and CODEAF_MUSIC_MODEL.

For one call only, name the model in the sentence: "draw this with gemini", "use seedance-2.0-mini", "the best video model". codeaf image takes --model the same way. A name that matches nothing is refused before anything is spent, and the result line always names the model that made the file.

Things to build

  • An ad campaign from one brief: a team with a copywriter chat, an image chat and a manager that reviews both, turning one product brief into three headlines, three banners, a 20 second voiceover and a music bed. Teams
  • Release notes read aloud: a watch on your tags that reads the newest CHANGELOG section into release-notes.mp3 each time a release is cut. Watches
  • A launch clip: a short product video and a jingle, cut together locally with the agent's video editing tool.
  • Docs diagrams that stay current: redraw an architecture picture whenever a module moves, feeding the last render back in as a reference.

Go further

Teams: Named, coloured groups of agents, gathered around the outcome they serve.

Coming soon: describe the org you want and CodeAF builds it: teams, managers, budgets, standing orders.