How to install the skill, plus a look at every image style and caption style it can produce.
Unzip it first — then install only the .skill file inside.
Full steps in Section 03 below.
Turns a topic into a finished narrated video: script, AI voiceover, styled 16:9 images, and an assembled MP4 with captions, a thumbnail, and upload-ready titles, description, and tags. Built for faceless channels, explainers, video essays, and documentary-style narration.
The approval gate: the skill stops for your sign-off on the script before it generates any voice or images. That's the cheapest place in the pipeline to change something, since voice and images cost real money and both get thrown away if paragraph two changes.
Used to plan beats, assemble the final MP4, and write captions. Usually already present in a Claude Cowork or Claude Code environment.
The image and (optionally) voiceover engine. Connect the Higgsfield MCP server
and the skill uses its generate_image_batch / generate_audio_batch
tools directly, no API key required.
ElevenLabs for the best voice quality, if you'd rather use it
instead of Higgsfield's voices. Needs your own API key in the
ELEVENLABS_API_KEY environment variable, never in config.json.
Images render at 2K, 16:9 so the camera moves stay sharp when cropped down to the final 1080p output.
Unzip before you install. The download is a zip archive with a
few files inside it. The one Claude actually needs is youtube-video-factory.skill.
Don't drag the whole zip, or the folder it unzips into, into Claude — open just that
one .skill file.
youtube-video-factory.skill is the file you need from it./youtube-video-factory..claude/skills/ folder.First run: if config.json is missing or incomplete,
the skill asks six short questions (voice provider, default video length, default
image style, niche, audience, channel name) and saves the answers so it never asks
again. Never store an API key in config.json itself, since that file
gets zipped and shared.
Ten built-in art directions, plus a slot for your own. Pick one as your default in
config.json, or choose per video.
Film-still realism. The safest default for serious topics.
Clean, bright, documentary-magazine photography.
Digital oil painting with visible brushwork. Warm and timeless.
Candlelit, aged, documentary-mystery. History and unexplained subjects.
Diagram-like cutaways and exploded views. For narration that explains mechanism.
Soft-lit 3D product-viz look. Modern, neutral, works for abstract ideas.
High-contrast graphic-novel panels. Bold and stylised.
Watercolour and sumi-e restraint. Elegant, calm, lots of negative space.
Feature-film anime backgrounds and key art. Strong for younger audiences.
Mid-century screenprint with a limited palette. Distinctive and cohesive.
Describe the style you want once. The skill saves it to your config so every future video can reuse it.
Every style deliberately avoids generated text in the image. Lettering from AI image models is almost always misspelled, and misspelled words on screen are the fastest way a video looks careless. Words go on screen as captions instead.
When burning captions into the picture, the look is fully configurable: font, size, bold, text color, outline color and width, drop shadow, and position. A few examples of what's possible:
DejaVu Sans, bold white, dark outline, bottom-anchored. What you get with no flags set.
Larger serif font, yellow fill, bottom-anchored with extra margin.
Small, light-weight, unobtrusive, positioned at the top of the frame.
Solid background box behind the text instead of an outline, for maximum legibility.