Back to Blog
September 30, 20267 min read

I Built a Machine That Turns My Articles Into Videos. Here's What Broke.

How I turned my articles into narrated videos without filming anything: a five-station pipeline on open models, the mistakes that cost me hours, and the one change that made the videos watchable.

Share:

Almost nothing that broke was the machine. I broke it: I described the stick man twice so the model drew two of him, I told the image model what not to draw instead of what to draw, and I once tried training a custom character model that ate the machine's memory and took the whole server down.

One night in August I made four videos back to back while I answered email. Each one ran under a minute. Each one was tied to a real California housing bill. Nobody filmed them, nobody edited them, and nobody sat in a voice booth. An article went in one end and a finished video came out the other, once widescreen for YouTube and once vertical for Reels and TikTok.

I run a brokerage, a property management company and an HOA management company. I write a lot. Legislation updates, landlord warnings, HOA explainers. The articles are solid, but text alone doesn't get read the way it used to. People scroll. They watch. The same idea reaches far more people as a video. And some weeks I don't have time to sit in front of a camera.

So the question was simple. What if every article I write could turn itself into a video, without me?

The stick guy

The videos have a host: a stick figure built to look like me. Same hair, same vibe, same guy in every video. Consistency is the point. Viewers start recognizing a character the way they'd recognize a person.

I gave him the hardest assignment I could find, my Sacramento Watch series on California housing and HOA legislation. Reserve funding bills and listing photo disclosure rules are about as dry as material gets. If he can make a reserve funding bill watchable, he can make anything watchable.

Rent the frontier, own the volume

My first version, back in February, rented everything. Every video had a meter running. So I flipped it. Open models on my own machine now do the volume work, and rented tools only handle the rare thing only they can do.

Five stations

Think of it as a production line.

  • 1. The script. I hand it an article or an outline. A local language model breaks it into scenes, one sentence at a time, and decides what the viewer should see during each sentence.
  • 2. The pictures. An open image model called FLUX draws every frame on a GPU box in my office I call Tim. The first image after install took about 32 seconds. After tuning, about 6.
  • 3. The voice. An open text-to-speech model reads the script. I pick the voice and the pace. My favorite runs about 12 percent faster than normal, because nobody wants a slow narrator.
  • 4. Assembly. ffmpeg, the free, decades-old video tool, stitches pictures, voice, captions and a music bed together. Twice. Once widescreen and once vertical, and the vertical one is a real vertical composition, not a crop.
  • 5. Delivery. Finished files land on a content delivery network with a YouTube title, description, chapter timeline and tags already written.
  • And the rule that matters most: it never publishes on its own. Every video stops at a preview gate. I watch it, I approve it, then it goes out. The automation does the work. I keep the judgment.

    Where it shows up

    On Townhome Pros, my HOA site, each Sacramento Watch article carries its own video right inside the post. The board member who won't read the whole article will watch a one-minute video. The widescreen cuts go up on YouTube as regular videos, the vertical cuts go out as Shorts, Reels and TikToks, and they land on TechTony.co too. One article, three or four places it can show up.

    The crew

    Nobody is paying me to name any of these.

  • Claude builds it. I didn't write this code. I directed Claude, the AI from Anthropic, like a developer on my team. When something breaks, it fixes it and writes the lesson into the playbook.
  • FLUX draws it. Every background, every scene.
  • Kokoro speaks it. An open voice model.
  • ffmpeg glues it.
  • The machines have names, because I'm that guy. Tim is the muscle, the GPU box that draws the pictures and does the voices. Clay is an always-on Mac mini assistant that can kick off jobs when I'm away from my desk. Bob is the backup that can run a job too.

    Full honesty: I still rent the fancy closed models for special projects, like the AI video clips in my apparel ads. You can't own those. And if a GPU box sounds like overkill, it is, for day one. Start rented. Once you know you'll make a lot of videos, that's when owning the machine pays for itself.

    The mistakes

    This is the part that saves you a week.

    Telling the AI what not to draw. I asked for a legislature hearing room with "no flags, no seals, no people." I got a flag, a state seal and a room full of people. Image models basically don't hear the word "no." Describe what should fill the space instead. "Bare wood paneling" beats "no seal" every time.

    Any writing surface turns into gibberish. Signs, documents, calendars, phone screens. The model invents words. I got "VACENT," "NOTGOOD" and a fake watermark. Bill numbers and dates now go in captions, never in the picture.

    Describing the character twice. The system already adds my stick man's look to every prompt. When I described him again in the shot, the model saw him twice and gave me giant floating heads and two men in suits. Short prompts, action only.

    The arrow. I asked for a crashed downward arrow. It pointed up. Twice. What finally worked was "hanging from the ceiling like a stalactite, arrowhead at the bottom." Sometimes you talk to AI like a very literal intern.

    The expensive one. I tried training a custom character model, and it ate all the memory on the machine and took the whole server down. Training now runs only inside a memory-capped container. Every lesson like that goes into a written playbook, so no AI agent, and no future me, makes the same mistake twice.

    The change that made them watchable

    The early videos were accurate, and they were boring. The picture didn't match what was being said.

    So I found a channel I actually watch and studied it frame by frame. One shot per sentence, and the shot is the sentence. A cut about every three seconds. Wide, then close-up, then an insert. A script that talks to you like it's your story, not a lecture.

    Then I built that into the machine: shot sizes, two-character scenes, props, and "holds," where it reuses the last background and just moves the character, so about half the shots don't need a new image at all. Same topic as an earlier version, 78 seconds, 32 shots. That's the difference between a slideshow and a video.

    It now has styles I can call by name. A truth-teller that opens with "here's the deal" and leans on ridiculous analogies. A countdown. A silent-film style for my apparel brand. Same engine, different personalities.

    Set up your own

    You don't need my setup. You need five decisions.

  • 1. Pick your rent versus own line. Start rented while you test. Move the volume to open models once you know you'll make a lot.
  • 2. Get one decent GPU box. That's what runs the images and the voice.
  • 3. Use an AI coding agent as your builder. I directed Claude, and it wrote the code, the tools and the fixes.
  • 4. Keep a playbook. Every expensive lesson gets written down so the next session starts smarter.
  • 5. Keep a human gate. Never auto-publish. Your name is on it.
  • This machine handles my faceless content. Some content needs a face, though. Mine. So I built a second system that takes footage of me, studies a creator's style I like, and edits my video to match. It edited the video that goes with this post. That's the next one.

    TT

    Tony Self

    AI strategist, speaker, and consultant helping enterprises deploy AI without the risk. Decades of experience in real estate and technology.