Much of the recent work has been programming the implementation of the radio drama interludes, done in the style of the old Starcraft 1 pre-mission scenes. While important, there was never any doubt that I could get something functional up and running, especially as the quality standards for a schlocky parody are not overly high. The harder problem is creating the content in the first place.
While good writing is difficult, my mock interviews with famous people are among my favourite works, and this is essentially an extension of that. Given enough time I can produce something that is at least satisfactory. The same cannot be said for the process of torturing some “AI” (Machine Learning) program into spitting out an acceptable performance.
It’s no secret that I’ve been extremely down on “AI,” partly for ethical issues, but mostly just because the results are terrible.

However, they’ve long proven themselves capable of at least moderate quality voice cloning. I believe these POL deepfakes date back to 2023, and the AI Presidents Tier List meme is from that same year.
I’ve even made a few of my own, with decent enough results upon a not particularly significant time investment.
However, one of the consistent problems with various machine learning products is that they excel at giving you something that is about 70% of the way to acceptable quality, and utterly fail at adjusting to your additional input, because they’re stupid bits of software that have absolutely no idea what’s going on. With voice generation software this results in an incredibly uneven delivery of lines. Some parts are exactly what you’d hope for, some are terrible, and the whole thing is inconsistent even with respect to basic features like volume, tempo, and intonation.
The celebrity parodies tend to land anyway, but that’s in large part because of the shock value in getting some milquetoast celebrity to speak like a 2015 anon. For copyright reasons, I probably can’t put any voice except my own into a game. Even if I could, the longer the content, the less funny the impersonation becomes, placing more emphasis on more traditional humour. However, good humour requires good delivery, and text to speech voice generation is kind of like hiring an actor who doesn’t speak English and can’t be given direction.
I recently struggled through these very issues implementing the announcer. To my great fortune, underneath that article I got a comment from “Chunk” who mentioned ElevenLabs’ voice to voice replacement feature. I gave it a shot, and while it has its own issues, with at lot of tweaking I got what you hear below.
The free voice I tried for my cohost sounded like a robot, but my voice came out pretty well. I coupled that with a free background audio track, added a few sound effects, and decided that the result was good enough, considering that it’s a quick and dirty placeholder clip. I’m a total amateur when it comes to audio editing, but I can already see that more production can overcome much of the weaknesses of this content. Enough that I’m going to spend the time creating a real attempt at the intro cutscene next.












