Every content creator we talked to said some version of the same thing: ideas are easy, production is hell. Recording a YouTube Short takes five minutes. The editing, the subtitles, the thumbnails, the actual publishing? That's hours. Someone shipping five videos a week spends most of their week in post-production, not in front of the camera. "I wish this wasn't so painful," one of them told us, and it stuck. So we started asking whether most of that pain could be handed off to software.
It could, mostly. But there's a long way between a clean idea and a product that holds up. "Use AI to find the best moments in a long video" sounds trivial until you build it. Our early versions cut subtitles off mid-sentence and grabbed blurry frames for thumbnails. Telling a model to "find the engaging moments" is easy. Teaching it what engaging even means is the hard part. Between audio analysis, facial expressions, and what's actually being said, it took months to get the different models cooperating.
These days, people using ShortsByAuto put out 15-20 videos a week. By hand they'd manage five, tops. For a small creator that three-to-four-times jump changes the whole math. One user put it plainly: "Before ShortsByAuto I was doing three videos a week and burning out. Now I publish fifteen and still have time left over."
The AI still isn't perfect, and we don't pretend otherwise. It picks an odd moment now and then, and the subtitles are sometimes wrong. But it does the first 80% and hands the last 20% back to the person, who's the one who should be making those calls anyway. That split works for where we are right now. We were never chasing perfect. We were chasing usable.