38 tools Avg score 8.4/10 37 with a free tier

AI tools for video and audio

The heavy end of generative AI. Credit-metered, render-queued, and licensing-sensitive in ways text tools never are.

Data last refreshed 2026-08-28

Where the money actually goes

Video and audio generation are the most expensive things in this index to run and the most expensive to get wrong. A rejected paragraph costs nothing. A rejected ten-second render costs credits, and the first week of using any of these tools is mostly rejected renders.

The list splits into generation and post-production. Generation is the headline: text to video, text to music, voice synthesis, avatar presenters. Post-production is quieter and often more useful day to day — stem separation, clip extraction from long footage, transcript-based editing, upscaling and cleanup. Teams that already shoot footage usually get more out of the second group.

Rights are a bigger deal here than anywhere else on the site. Training-data provenance, voice cloning consent, music licensing for commercial release, likeness rights on avatars. These are not edge cases, they are the first question a client legal team will ask.

What carries this tag

  • Text-to-video and image-to-video generation, including avatar presenters.
  • Music generation, voice synthesis and cloning.
  • Post-production: stem separation, transcript editing, clip extraction, cleanup.
  • Image tools appear only where they feed a video or audio pipeline.

Which categories these come from

34 of 38 are free or freemium, 4 are paid only.

The top 12, ranked

Ranked on AI Score, then adjusted for how recently the tool shipped, whether it is an editor's pick, and what readers actually open — so a strong recent release can edge out a slightly higher score. Every entry shows what it is good for and what it costs you, including the parts the vendor leads away from.

1 Runway logo
Runway 9.3/10 Free tier + Standard $12/mo + Pro $28/mo + Max $76/mo + Enterprise contact sales ★ Pick Video

AI video platform with Gen-3 Alpha. Create cinematic video from text or images with realistic motion, lighting, and physics. The go-to tool for creative professionals.

Best for Filmmakers, advertisers, and content creators who need cinematic AI video with realistic motion and fine-grained creative control.

  • Motion Brush lets you paint motion paths onto video
  • Comprehensive creative suite beyond generation, including a full video editor
  • Gen-3 Alpha model for temporal consistency and physics-accurate scenes
Free tier

125 credits (one time)

Trade-offs
  • Credit system can get expensive for heavy use
  • Short clip duration (5-10 seconds)
2 Suno AI logo
Suno AI 9.2/10 Free tier + Pro $8/mo + Premier $24/mo ★ Pick Music

Generate full songs with vocals, lyrics, and instrumentals from a simple text prompt. Supports dozens of genres and styles, from pop hits to lo-fi beats to orchestral scores.

Best for Content creators, podcasters, and indie artists who need original full songs with vocals and lyrics without studio production.

  • Generates complete song structures - intro, verses, chorus, bridge, outro - not just loops
  • Covers virtually every genre, from pop and jazz to lo-fi and orchestral
  • Commercial license available on paid tiers, free tier is non-commercial only
  • Free tier allows 10 songs per day for experimentation
Free tier

50 credits renew daily (10 songs), no commercial use

Trade-offs
  • Non-commercial license on free tier
  • Limited control over specific sections
3 ElevenLabs logo
ElevenLabs 9.2/10 Free $0 (10k credits) + Starter $6/mo (30k) + Creator $22/mo (121k, first month $11) + Pro $99/mo (600k) + Scale $299/mo (1.8M) + Business $990/mo (6M) + Enterprise custom. Credits are shared across products. Annual is 10 months prepaid. ★ Pick Music

AI voice platform for text-to-speech, cloning, and agents. Hosted MCP lets Claude and Claude Code manage those agents over OAuth.

Best for Creators, publishers, and developers who need realistic TTS or cloning, and teams that want Claude or Claude Code to manage ElevenLabs agents over hosted OAuth.

  • Voice cloning from a short sample, including professional clones on Creator and up
  • One shared monthly credit pool across speech, music, effects, and dubbing
  • Hosted MCP with OAuth in Claude and Claude Code — no API key in the client
  • API with streaming for real-time apps, plus Studio for long-form
Free tier

10,000 credits per month, non-commercial, no rollover. Shared across products. elevenlabs.io/pricing 28 Aug 2026.

Trade-offs
  • Credits are shared, so music or dubbing will eat the TTS budget faster than the headline number suggests
  • Free is non-commercial and does not roll over unused credits
4 ChatGPT logo
ChatGPT 9.5/10 Free tier + Plus $20/mo + Pro $200/mo ★ Pick Chatbots

OpenAI's flagship AI assistant with o3/o4-mini reasoning, GPT-4o, Advanced Voice, Sora video gen, Operator agent, and Deep Research — the most feature-packed chatbot available.

Best for Users who want one subscription covering reasoning, voice, vision, image and video generation, and agentic browsing in a single app.

  • Broadest feature set: voice, vision, image gen, Sora video, browsing, and agents together
  • o3 and o4-mini reasoning models for complex math, science, and coding
  • Operator agent and Deep Research add autonomous multi-step task completion
  • Free tier now includes limited GPT-4o access, not just the mini model
Free tier

Free access to GPT-4o mini and limited GPT-4o with basic features

Trade-offs
  • Pro plan at $200/mo is hard to justify unless you need heavy o3 or Operator usage
  • Writing quality has fallen behind Claude for nuanced, long-form content
5 Veo 3 logo
Veo 3 9.1/10 Free via Gemini + Vertex AI pay-per-use ★ Pick Video

Google DeepMind's flagship video model generates cinematic clips with synchronized native audio from text prompts.

Best for Filmmakers, content creators, and marketing teams who need production-quality cinematic video with synced audio, without a production budget.

  • Native synchronized audio (dialogue, sound effects, ambience) generated with the video
  • Cinematic camera, lens, and lighting control via natural-language prompts
  • 4K output with strong physical consistency for water, fabric, and hair
  • Accessed through Gemini Advanced or Vertex AI rather than a standalone app
Free tier

Limited generations available through Gemini with a Google account

Trade-offs
  • Locked into Google's ecosystem with no standalone app or open weights
  • Vertex AI pricing can add up quickly for high-volume production use
6 Krea AI logo
Krea AI 8.6/10 Free tier + Basic $9/mo + Pro $35/mo + Max $105/mo + Business $200/mo + Enterprise Custom ★ Pick Image Gen

A real-time creative platform that generates and refines AI images, videos, and 3D assets interactively as you type and sketch.

Best for Designers and creative teams who want to iterate on images, video, and 3D assets live on a canvas instead of queuing single prompt batches.

  • Real-time canvas updates as you type prompts, sketch, or drag reference images
  • Combines image, video, 3D, and enhancement tools (upscaling, background removal, style transfer) in one platform
  • Free tier lets you test the full workflow before paying
  • Faster iteration than prompt-and-wait tools, at the cost of some peak image quality
Free tier

$0 /month. 100 compute units / day. Access to Krea 2. No credit card required. Full access to real-time models. Limited access to image, video, 3D, and lipsync models. Limited access to image upscaling. Limited access to LoRA training.

Trade-offs
  • Peak image quality slightly below Midjourney for final artwork
  • Video and 3D features are still maturing compared to dedicated tools
7 Seedance 2.0 logo
Seedance 2.0 9/10 Free daily credits (Dreamina) + paid ~$15-$70/mo; API from ~$0.08/s ★ Pick Video

ByteDance's multimodal video model generating up to 15s of 1080p multi-shot footage with native synced audio, lip-sync, and director-level camera control.

Best for Video creators and developers who need short AI-generated clips with synchronized audio, lip-sync, and cinematic camera control.

  • Generates dialogue, effects, and lip-synced audio alongside the video, not dubbed on after
  • Ranked #1 on the Artificial Analysis Video Arena for text-to-video and image-to-video
  • Director-style camera controls with multi-shot consistency across cuts
  • Accepts text, image, audio, or video as conditioning inputs
Free tier

Free daily credits via Dreamina with no card required (~2-3 short clips per day)

Trade-offs
  • Capped at 15 seconds per generation — short of what longer-form projects need
  • No native Dreamina API; programmatic access depends on third parties like AtlasCloud and PiAPI
8 Luma Dream Machine logo
Luma Dream Machine 8.8/10 Freemium — paid from $29/mo ★ Pick Video

High-quality cinematic video generation from text or images using Luma's Ray2 model, known for realistic physics and fast rendering.

Best for Filmmakers, advertisers, and concept artists who need fast cinematic pre-visualization or storyboard clips from text or images.

  • Realistic physics for water, fabric, smoke, and particle effects
  • Generation times faster than most competitors
  • Strong image-to-video pipeline that preserves reference-frame style for storyboarding
Free tier

Limited free daily generations with watermark — enough to evaluate quality

Trade-offs
  • Complex multi-character scenes can lose coherence
  • Clip duration is shorter than Kling AI's longer-form output
9 HeyGen logo
HeyGen 8.8/10 Free tier + Creator $29/mo ★ Pick Video

AI video platform that turns text scripts into realistic avatar videos with lip-sync, multilingual translation, and personalized video at scale.

Best for Marketing and sales teams producing talking-head avatar videos, multilingual dubs, or personalized outreach clips at scale.

  • Video Translate clones the speaker's voice and re-syncs lips across 40+ languages
  • Personalized Video API generates thousands of individualized outreach videos from one template
  • Interactive Avatar supports real-time conversational use cases like AI receptionists
Free tier

Free tier lets you test avatar creation with 1 credit — enough to evaluate quality but too limited for real production

Trade-offs
  • Free tier is extremely limited — essentially a demo with 1 credit
  • Stock avatar quality trails Synthesia's top-tier options slightly
10 Chatterbox logo
Chatterbox 8.8/10 Free MIT open-source model + paid Resemble AI hosted platform ★ Pick Music

Resemble AI's open-source (MIT) text-to-speech and zero-shot voice cloning model with emotion control, 23+ languages, and a watermark on every output.

Best for Developers and teams who want to self-host an open-source, zero-shot voice cloning and text-to-speech model instead of a closed API.

  • Fully MIT-licensed open weights: free commercial use, self-hosting, no royalties or caps
  • Zero-shot voice cloning from a short reference clip, no per-voice training run
  • Emotion and intensity control plus 23+ language coverage via the Multilingual line
  • Imperceptible neural watermark on every output by default, unusual for an open-weights release
Free tier

The complete Chatterbox model is free under the MIT license — self-host it, use it commercially, no usage caps. You only pay if you opt into Resemble AI's hosted platform.

Trade-offs
  • Self-hosting needs your own GPU and technical setup — there's no polished consumer app for the free model
  • Quality and naturalness vary across the 23+ languages; English is the strongest
11 Descript logo
Descript 8.7/10 Freemium — paid from $8/mo ★ Pick Video

AI-powered video and podcast editor that lets you cut footage by editing a text transcript, with Overdub voice cloning and eye contact correction.

Best for Podcasters, YouTubers, and internal comms teams who edit talking-head or interview footage and prefer working from a transcript.

  • Cuts video by deleting words from the transcript instead of scrubbing a timeline
  • Overdub clones your voice to fix flubbed lines by typing new speech
  • Eye Contact correction and one-click Studio Sound cleanup replace manual post-production steps
Free tier

One project with watermark — enough to test the transcript-editing workflow

Trade-offs
  • Not suited for complex VFX, motion graphics, or color grading
  • Overdub voice quality can sound slightly synthetic in longer passages
12 Viggle AI logo
Viggle AI 8.5/10 Free tier + paid from $9.99/mo ★ Pick Video

AI character animation tool that makes any person or character move realistically from a single photo — no animation skills required.

Best for Meme creators and social video makers who want to animate a single photo of any person or character without animation skills.

  • Physics-aware motion from one static image — cloth, hair, and weight shift naturally
  • Mix mode transfers motion from a reference video onto any character image
  • Large, constantly updated library of viral dance and motion templates
  • Purpose-built for meme/social exports rather than cinematic scene generation
Free tier

Daily credits with watermarked output

Trade-offs
  • Free tier output is watermarked and resolution-limited
  • Best at human-shaped characters — non-humanoid subjects often struggle

The other 26 in the index

Same tag, lower down the ranking. Scores, pricing and full write-ups behind each name.

Opus Clip Video · Freemium — paid from $19/mo 8.6/10 Kling AI Video · Freemium 8.9/10 Udio Music · Freemium 8.8/10 Synthesia Video · From $22/mo 8.7/10 Leonardo.ai Image Gen · Free tier + from $12/mo 8.7/10 Muse Image Image Gen · Free in Meta AI + paid Meta subscription for higher limits 8.5/10 Pika Video · Freemium 8.5/10 AIVA Music · Freemium 8.4/10 Grok Imagine Quality Mode Image Gen · API-based pricing (check console.x.ai for current rates) 8.4/10 Stable Audio Music · Freemium 8.4/10 Creatify AI Video · From $29/mo (credit-based) 8.3/10 LALAL.AI Music · From $15 8.3/10 CapCut Director Mode Video · Free with 300+ AI credits + Pro tiers 8.2/10 ZONOS2 Music · Free (open-source) + paid cloud tiers 8.2/10 Beatoven.ai Music · From $14/mo 8.1/10 Vidu AI Video · Free credits + paid plans (check site for current tiers) 8/10 Mubert Music · Freemium 8/10 Higgsfield AI Video · Free trial credits + paid from ~$15/mo (Starter) to ~$84/mo (Ultra) 8/10 Loudly Music · Free tier + from $9.99/mo 8/10 Magnific Image Gen · Free tier + paid from ~$5.75/mo (credit-based) 8/10 Captions (by Mirage) Video · Free tier + Pro $9.99/mo, Max $24.99/mo, Scale $69.99/mo 8/10 Tavus Video · Free tier with credits; usage-based paid plans + enterprise (check site for current rates) 8/10 ElevenMusic Music · Free tier + Pro $9.99/mo 8/10 Varya Video · From ₹0.48/sec (~$0.006/sec) 7.4/10 Kits AI Music · Free tier + paid from ~$10/mo 7/10 Masterchannel Music · Paid plans (EUR ~€180–€948/yr); reviews cite ~$15–$20/mo annual — check website for current pricing 7/10

Open all 38 in the filterable directory →

Where to start without paying

37 of the 38 tools here publish a free tier. These are the strongest of them, with what the free tier actually gets you.

Runway Free tier

125 credits (one time)

Suno AI Free tier

50 credits renew daily (10 songs), no commercial use

ElevenLabs Free tier

10,000 credits per month, non-commercial, no rollover. Shared across products. elevenlabs.io/pricing 28 Aug 2026.

ChatGPT Free tier

Free access to GPT-4o mini and limited GPT-4o with basic features

Veo 3 Free tier

Limited generations available through Gemini with a Google account

Krea AI Free tier

$0 /month. 100 compute units / day. Access to Krea 2. No credit card required. Full access to real-time models. Limited access to image, video, 3D, and lipsync models. Limited access to image upscaling. Limited access to LoRA training.

See all 37 free and freemium options →

What decides it

Clip length and resolution ceiling

The published maximum is the number to design around. Stitching short generations into something longer is where consistency falls apart, and no amount of prompting fixes a hard cap.

Credit arithmetic at your reject rate

Work out the cost of one usable output, not one generation. On generative video the ratio between attempts and keepers is the whole budget.

Rights and provenance

Commercial release terms, training-data disclosure, and whether consent is verified for cloned voices. If the vendor is vague, assume the answer is one you would not like.

Does it hand back editable pieces

Separate stems, per-shot files, project export. Tools that only return a flattened final render force a full regeneration for any note, which is fine for a demo and painful in production.

What still does not work

  • Character and scene consistency across shots remains the defining weakness of generative video. Every vendor is working on it and none of them have solved it.
  • Lip sync, timing and audio alignment degrade the further you get from short, front-facing, single-speaker footage.
  • Long-form output is a stitching exercise, not a generation one. Plan for an editor in the loop, because the tool is not going to be one.
Comparison explorer Put the top three side by side Opens with Runway, Suno AI, ElevenLabs already loaded. Swap any of them out and read pricing plans, features, pros and cons in one table.

Related reading

Browse another job

Weekly issue

The 5 AI tools that mattered this week.

One email, Fridays. No spam, unsubscribe anytime.