# How Claude Opus 5.5 Can Run a 24.5k/Mo Faceless Youtube Channel Channel: g russ Video: https://www.youtube.com/watch?v=8m6-EmYW2zQ Duration: 28 min Language: English Words: 5038 Transcript page: https://viewrankai.com/tools/youtube-transcript/8m6-EmYW2zQ --- [0:00] Today I'm going to show you how to go from zero to three finished low poly YouTube shorts made with Claude Opus 5.5. The script, the voice, the animation, and the little tricks that make the whole thing look extremely good without filming yourself or spending hundreds of dollars on a team of animators. And then we are going deeper because getting Claude to make a video is one thing. Getting it to make something that you actually want to post is where it gets interesting. But you're probably thinking to yourself, "Oh, I've seen so many of these faceless AI [0:28] YouTube videos and gurus and whatnot." You paste a script, you generate a voice, and end up with a generic ass video like everybody else. And look, I get that. You see, we are going to start by actually breaking down what works in the channels that we are studying. How they hook you, how they tell you the story, and what makes you want to see how it ends. Then we are going to use [0:46] that formula to do something original. And from there, I'm going to show you how to give it a voice and have Claude build the whole 3D scene around it which matches the actual spoken words. So when the story moves, the visuals move with it. And this is the part which I really want you to see. We are going to make Claude look at its own work, fix what's wrong, and then use visual references to improve the lighting, the textures, the [1:08] details, and the whole cinematography. So, I'm going to walk you through every single step, show you the prompt, the early versions, final result, and the ultra 3D skill, and explain how to repeat the process for your next video. And we are going to do all of that with Opus 5.5 at low effort, which is actually crazy. And if you are still doubting me, this is my actual YouTube play button that I got with my main channel on which I made animations with AI. Okay, let's jump in. All right, hope you like that intro. That was actually edited with my ultra edit skill inside cloud right here. as you can see. But [1:40] today we are actually talking about how Cloud can run a 24.5K per month faceless YouTube channel. Now I have the whole project right here inside my cloud code. We have this PlayStation one style project. If I scroll up to the top, you can see that it can't even load the first messages. But don't worry, I got you covered. I made this whole 39 slide presentation on the whole actual workflow along with three exact examples that we are going to take a look at and show you everything step by step. So the two channels that we are taking a look at today as a reference, one of them is low [2:16] poly shorts and the other one is loaded dice probably has the same owner because of the same handle but nonetheless they have a really unique and distinctive short style that is extremely hard to replicate. diffusion based models, even cedance 2.5 is struggling with this type of style. And even if you use blender pre visualization, it's kind of hard to get this exact style right. And so what Opus actually allows us to do right here is use something called 3JS as the rendering engine and to build on top of it, place the assets, place the camera movement, everything inside of it and create a finished render. Now inside 3JS [2:57] it can also use 3D models that you can either give it from the internet or you can make your own. It can create the textures for it, rig the characters and you have essentially full control over it. And I would say it's easier for beginners with zero experience to iterate on this one than to use Blender. Now on that note, low poly shorts is bringing in half a billion views in the last 48 shorts they made. On average, they get 7.4 million views, and that's across both channels. The reason why these channels blowing up so hard is because this PlayStation one low style is something that you can instantly spot [3:36] in your feed, and it's so distinctive because it's insanely hard to make. Arguably, it's harder to achieve this aesthetic than the Zagd film style, but that doesn't matter. We have Oppus 5.5 now. So, if you spend a considerable amount of time training it properly, which I'm going to show you in a minute, then you can get some pretty nice results. Of course, it's not perfect, but it's 90% there. And for mass production, that's what matters. Because the thing that you have to notice is [4:03] that this channel only has 85 uploads. And that managed to get them 1.2 million views. And if you look at their upload cadence, they are uploading every 2 weeks. With this cloth scale that I'm giving you, you can do this every single day like 10, 20, 30 times, and it doesn't even cost that much credits. Yes, I'm going to cover the credits as well, but honestly, overall, it's pretty cheap. Now, just as a quick reference, I'm playing a little clip from one of their videos, and honestly, props to these guys because the quality, the sound design, and the whole aesthetic is just so unique. But we are trying our [4:38] best to mimic this and create something unique. So here's what we're going to build. It's going to be a complete system where we have the script. Based on the script, we use GR text to speech to generate the voice and we give the voice to Claude. Claude is going to manage 3JS, set up everything, get a rich hero, the cameras, the captions, and it's also going to autoche its own work. Then we are going to have a finished short in 9x6 4K resolution. And this can be anywhere between 30 seconds to 90 seconds. whatever you like. Now, for this video, I made three examples which I'm also posting on my main [5:16] channel and all of those three are backrooms shorts. So, version zero is going to be from scratch. We don't give Claude any references. We only give it the voice and we let it figure it out on its own. Then, in version one, it's going to be with references from either low poly shorts or you can just simply go to Pinterest and put in like five images. That's what I did. And version two and version three are essentially selfiteration loops where it's going to visually verify its work and fix any inconsistencies because our goal is to create vertical educational shorts in a PlayStation one look as close as [5:52] possible to low poly shorts. So here's the actual three versions side by side. In 2019, someone posted a photo of an empty yellow office online and a reply said that if you glitch out of reality, this is where you land. They called it the back rooms. You see, nobody knew where the picture came from. No names, no location, just damp carpet, yellow wallpaper, and buzzing lights. So, for 5 years, people argued it was a render, a [6:17] dream, or a place that never existed. Then, in 2024, a group of internet detectives dug through old archived websites and found it on a hobby store's renovation blog from 2003. The scariest room on the internet was a furniture store halfway through becoming a hobby shop. All right. So, as you can see, without references, only providing audio and saying, "Make me a 3D animation really using 3JS didn't really make [6:45] something that's usable or postable." Version 3 by any means with the references that you provide is already kind of usable. Matter of fact, it's good. But version three actually managed to take it up a notch. Multiple scenes like this one looks a lot better. It's more polished here as well. It's more detailed. Character animation looks different. The screen cinematography is just a little bit better. And here as well, we have just more details. So overall, it's just stacking details upon details to make the whole thing look more polished. So here's actually changed between those versions. In version zero, we did everything from scratch. There were no references, no [7:24] assets. It only used boxes and spheres. We had a boxman. And Cloud actually never verified its work. it just went ahead and rendered. Now at version one with references which we can provide from Pinterest or from our reference channel, it managed to nail the style completely and it had that PlayStation one aesthetic that we want to achieve. Now with version three, it went through multiple self iteration loops and I'm purposefully not putting version two here because it's kind of the same. image just had multiple renders until it managed to fix the little things that it found as problematic because references get you most of the way but self iteration where [8:03] it actually visually verifying its work is going to give you the best results. Now there's five easy steps to automate this whole thing. Step one is doing the research. So studying the references, getting the transcripts from the videos, understanding how the actual story arc is being built and then reverse engineering that whole concept so that Cloud can make itself a skill internally that it can use to write similar scripts in the tonality which are completely unique. So we are not stealing anything, we are learning from it. And then at step two, once we write the new scripts, we actually want to humanize them. And there's a really handy tool for that [8:41] called the humanizer skill on GitHub. Let me show you that. It's from this guy called Bladder. And he has this humanizer skill. So you literally just copy this, give it to Claude, and say, "Bro, please humanize it." And he's done. Because this thing is going to get rid of a lot of language patterns that AI has by default. Now, step three is actually making the audio. Instead of recording, you can use a texttospech [9:06] engine. For example, I like using rock. You go to console.x.ai. You can choose any voice here. It's completely free. Select your speed 1.1 studio quality and you generate the script. Simple as that. Now, step four, building your version one. This is where the PlayStation engine gets built. The hero is getting rigged and the cuts happen based on the transcript. And there's a quality audition gate. And in step five, it's actually going to look [9:34] at its work, review it, and improve it. self iterate constantly. That's where the ultra 3D comes in. But you might ask, why use low effort when we have medium, high, extra high, even ultra code? The thing is, I've seen a lot of these guys online on YouTube putting editing agents and everything on ultra code. And I wait for 3 hours for a video to get edited only for them to see something they don't like a little text or some kind of sloppy motion, a zoom in or something they want to change and it takes like 30 to 40 minutes for a quick iteration when in reality it should take [10:14] like 2 minutes really. If you use low effort that doesn't mean that the intelligence gets cut in half. is just trying to do the work as quickly as possible and it's not paying as much attention. But if you purposefully instruct it to have a quality standard and inspect its own work like I did with my Ultra 3D skill, then it's going to put in high effort while working faster and burning less credits because the whole animation work is a loop. So you want to write a scene, you want to render it out, we want to look at it and fix anything that's wrong. And we are [10:48] going to do that over and over again until we get the style that we are striving for. You want to make each of those turns fast as possible and cheap. Of course, low offer just simply means less thinking per turn. So the replies come back quicker and burn fewer tokens. Quality comes from the loop, not from one long thing. So what does it actually cost? Well, it says $4 million per token in and $20 out, and it's 20 cents to read from cash. Now, a render loop rates the same project files constantly. So, most of it comes out of the cache, which is, by the way, a ridiculous price. I [11:28] had cloud summarize this project. In total, I burned 1 billion tokens. But here's the crazy part. I'm on the 5x plan right now and on my weekly limits I used roughly 60%. Now when I started to work on this project I was at 30%. So on the 5x plan it took 30% of my weekly plan. Do whatever you want with that information. If you are on the 20x plan you won't run into any issues and the whole token consumption in this case can be insanely optimized once you don't [12:02] reread from the same file and cache. When you have the whole skill logged in and base informations so that it only needs to write the code and do the iteration, it's going to be a whole lot more efficient. In this example, it says that the cash reads would have costed $500 roughly if it was charged on the API, but we are on the subscription, of course. Cash rights is $66 and the output is $49. But again, I'm on the 90 plan and it only used 30% of my weekly limits. So this is a bit far from the reality. But overall, given the fact that I built the whole rendering engine [12:44] from scratch, run it through multiple iterations way before I started giving it references, I would say I'm pretty happy with this because after this, it only gets cheaper. How do you steal the successful formula? Well, you have to study what already works before writing anything. So, if you go and study the 10 million view short, then you're going to understand how to structure it. We are not trying to copy from them. We are trying to learn the pacing and the patterns. So, you can go to the competitor's YouTube channel, sort by popular, grab one of the videos, copy the link, go to something like YouTube transcript, boom, paste it, and now you [13:24] have the full transcript. You can give this to Claude and it can analyze it. Now you can do this with multiple shorts or you can also automate the whole process of getting the transcripts of every single video. But the point is that mixing the topics and the transcripts is completely fine for CL to understand because we are trying to study the structure not the subject. We are trying to understand what makes those videos so great without copying them. Now the low poly shorts formula looks something like this. Based on the 83 shorts that I managed to analyze. On [13:55] average they are 34 seconds long. There's 115 words for each script. 3.4 words set per second. Every 2.3 seconds there's a hard cut. And those are the main things Cloud needs to understand before it starts working on any of those videos. Because if the pacing is off, if there's not enough hard cuts, if the text to speech is slow, then people are not going to stick around. They are just going to swipe. You're going to have a terrible swipe through rate and the [14:25] completion rate would be even worse. Now, there are also certain elements to this like a recurring everyman character. Every single video has the same character. It has literal visual gags, small yellow captions, and an ironic last line. So you can transcribe those shorts as cloud for the same like this and then you can use those numbers to prompt claude as hard rules. Now the actual analysis can look something like this. You can simply go ahead screenshot this grab this give it to Claude and you can begin your analysis. And once you're done with the analysis and it knows the transcripts and the patterns, then we use the second prompt to have new ideas [15:06] and to write new scripts based on the parameters that I described earlier. After we add the copy, we should make it sound human. So as I talked about the humanizer skill, this is the part where we actually go ahead grab this paste it into cloud after the scripts and it's going to run that humanizer press. Now this is incredibly important. We want to skip any kind of funky AI generated sloppy text. So as a quick recap, you go ahead and open low poly shorts on the shorts tab sorted by popular. You copy three links into YouTube to transcript.com. You open a new cloud chat oppus 5.5. Paste prom 0.1 with the [15:49] transcripts. Show the formula it writes and point at the quotes. Then you paste prompt 02. pick an idea and get some scripts. Then lastly, you run the humanized air pass and you are done. Okay, chapter two, generating the actual voice. So, as I said earlier, I like using Gro voice. So, you can just simply go to console.x.ai, choose voice texttospech. That's it. You simply paste the humanized script there, pick any voice you like, set the speed [16:20] to 1.1x, generate and download the MP3. You can paste this into your project folder. Now, quick side note. The reason why you want 1.1 instead of 1x short reward pacing and fast action and a slightly quicker voice feels like someone is telling you a story they can't wait to finish. And also, we are going to have a higher average view duration and watch completion, which is pretty important. Now, one last thing we do is cutting out the silences and not the words. What I mean by that is once we have the AI generated audio, there's going to be gaps between the spoken words. So, we want to give Claude a [17:03] prompt like this, which is going to tighten up the whole thing. And because of that, in my example, shorts, the length can go down by a couple of seconds, by up to 5 seconds, which is going to make the whole real feel much much snappier. So, again, here's a quick recap. You go to console.x.ai AI open text to speech. You paste your script and then then you select one voice that [17:28] you really like. Set the speed to 1.1x. Generate it. Listen to it that you actually like it. Then you download it. Put it into your project folder. You can rename it to voice.mpp3. Chapter 3. Building your version one. Okay. So you make a new folder. You can call it whatever you want. Then you drop in your voice.mpp3 and your script. TXT. Then you open cloud code. Click new session and choose that folder that you just created. Now in the model menu you can [17:56] pick opus 5.5. Insert it to low effort. Then you are going to paste the build prompt and send it. This is the build prompt and cloud is going to set you up the 3JS for the 3D scenes. A headless browser to render every single frame ffmpeg along with whisper for the video and the frame perfect timestamps so that it can create each scene. This is the part where you show it what good actually looks like. So you can save up to five to 10 reference images in a folder called references. And Claude is going to go see that and try to rebuild those low poly models and build a whole [18:32] engine around it. That's the difference between version zero and version one because version zero had nothing to look at. Yesterday night before I went to sleep, I actually told Claude that, "Hey, I have all of these res, I have all of these images, all of these references in this folder, and I want you to create YouTube shorts for me with audio based on those references. And along the way, go ahead and document every single mistake, everything you might have learned because I want to have those things all logged in." And because of this, I managed to build my own engine which I used for the references to get the results that you [19:09] saw at the beginning of this video. One more time, here's the version one prompt. You can pause the video, take a screenshot, get into your code. Now, here's where the actual PlayStation one look comes from. It is one custom shader inside 3JS that does the PlayStation one part. So, every limit of the old [19:28] hardware is faked on purpose inside 3JS. For that, we got vertex snapping, which is going to snap the corners to the coarse grid. So, the models are going to wobble like the real PlayStation one polygons. There's also a fine textures where the textures are going to warp across big floors and walls, which causes the famous PlayStation one swim. It also has guard lighting, fog, dither, and five bit color. All of these things made up the original PlayStation 1 aesthetic as we know it. Now we have an internal resolution of 540 * 960 and later when we are rendering that's upscaled to 4K. The headless renderer runs at 23 fps and it pipes into ffmpeg [20:12] and one 4K render takes roughly 45 seconds and that's on my RTX 5080. Now here's how the editing and pacing boils down. When we give Claude the actual voice text to speech, it's going to use a transcription model like Whisper to have a frame perfect understanding of where each word lands. And because of that, Claude knows exactly where to cut those scenes, how to picture things as the voice is moving along. And because of these word level timestamps, it can actually understand the scene a lot better than just simply going from the script by itself. convert level timing for building the whole thing is the most [20:53] crucial thing. And then I gave it a hero. So low poly shorts has the same character in every single video. Our version zero had a box character with a sphere on top. And by the time we arrived to version three, we had a real rigged PlayStation X style character with a 65 bone skeleton custom pose layer on top of it where the limbs were actually moving and it was actually animated. And now again this is why [21:18] every check matters before the render. So before any final render an automatic check looks at every bit and error blocks the render. We are looking for a quality standard. So this prompt helps us build that that standard by looking for subjects either in the wall cropped camera inside blown out highlights any visual inconsistency and it's going to [21:40] mark those down so later it can fix it. So once again here's a quick recap. First you have one folder where you put your text to speech along with the script. Then inside a new class session, you pick that folder. Put Opus 5.5 on low. You paste your prompts and let it plan. As it's planning and building the word engine, it's also going to work on trying to recreate that PlayStation one aesthetic, rendering the scenes and fixing the steels. Then you are going to have your version one ready. Now, as I said at the beginning of this video, I have two more examples. So, let's take a [22:15] look at them. If you've ever walked through an empty mall at night and felt your skin crawl, you already understand the back rooms. You see, hallways, waiting rooms, and office floors are built to be passed through by crowds. Your brain has seen them full of people thousands of times. So, when they're empty, it keeps expecting someone to walk around the corner. The internet calls these liinal spaces from the Latin word for threshold. Then add old fluorescent tubes. They hum and flicker 120 times a second just below what you consciously notice. The back rooms takes all of it and removes the exits. So the scariest thing about the back rooms [22:50] isn't a monster. It's that nobody is there. I think you can understand how crazy this is actually because the implications of this are incredible. Using 3GS as a rendering engine for these animations allows you to have insanely quick iteration. And if you build out the whole engine just like I did, then you can just put in multiple assets, actual characters, rig the animations, have more details. You can build a whole production pipeline around this. So here's the last part, reviewing it. It's really important for Cloud to actually review its work. If you set the effort to medium, high, or even max, it's going to do it by default usually, [23:28] but it's going to burn a lot more credits. is going to spend time that doesn't equal to actual quality improvements. That's the reason why we use a separate reviewer that only sees the script and the frames of the render. So, it doesn't know anything about the actual engine and what went into it. It [23:45] doesn't know how it's supposed to look. It just knows what it looks like. So, it can judge its work and go back to claw and tell, "Hey, bro, this this is way off." Now, as you can see, most of the improvements come from round one into round two. So, that's why we have version one, two, three. So, it can go [24:04] ahead run those issues and fix them. Now, I'm actually going to show you a couple of real life examples that this audition that this review process managed to catch. Right here, we had in version one only the head was glitching through. We didn't had the actual body. And then at version three, the whole body glitched through. In the third short, which I'm going to show in a minute, we had the main character in a really weird pose floating, not even being on the ground, trying to glitch through the wall. I'm guessing in version three, it actually there's an actual wall and the character walked into it. Now, from short two, you [24:39] probably seen this before. The brain didn't really look like a brain. It looks something else. But it's fine cuz the review caught it. And now we have something that looks like a brain. So good job cloud. Here are some small things that actually matter. So in short one, we had the 2003 blog post and the store name was cropped. So the whole so the composition was kind of bad. By version three, it got fixed and and also in short one at the end of the video at the payoff in version one it was kind of boring. By version three we had way more details, a leading line, new [25:17] composition, and a couple of elements that just tied the whole short together. And that's exactly what the Ultra 3D skill does. That's what I used to have version 1, two, and three with this whole self iteration loop built into it along with the engine that I trained overnight. It takes care of all the assets from the references, the lighting, and the camera. It does all the checking, renders, it reviews it, and fixes everything before you even see the final result. Now, here's the third and last short, which I want to show you. If you type a secret, if you type a secret code into the original Doom, you [25:49] can walk straight through walls. Gamers call it no clipping, and it's exactly how the back room says you get in. You see, walls in old games are paper thin shells painted on one side only. Step through one, and there's nothing behind it. No floor, no ceiling, just an endless void because the game only draws what it expects you to see. So, the back rooms asks a creepy question. What if reality worked the same way? What if behind every wall you'll never walk through, there's a place nobody finished building, and if you ever end up there, the carpet is wet, the lights are buzzing, and there's no cheat code to [26:22] get back out. That might actually be my favorite short. And again, just because you owe it the quality, that doesn't actually mean that you pass the quality standard because checks cash short you can describe in advance, but self iteration catch is catching the rest. That's why you measure the reference, automate what you can, let it review itself, and log every failure as a rule so next time when Claude encounters it knows exactly what to do with it. And then you just go and do the whole thing over and over and over again. You pick the winners, you batch the script, you batch the voice, you batch the TTS, put [26:58] everything into one folder, and then you post it. Now, this whole system can run on autopilot and you can access my build with all the files, all the assets, skills, rendering engine built out completely inside my school community. Right now, I'm running a pre-sale at $29 a month before I finish all the materials inside, then it's going to go to $97 a month. So, if you're interested in this workflow that I just broke down and you want to get constant update and and lock in your price for lifetime, make sure to check the first link in the [27:30] description. Thank you for watching. --- About this transcript Read from YouTube's own caption track and laid out by ViewRank AI (https://viewrankai.com). ViewRank AI finds the videos already beating a creator's own average on Instagram, TikTok and YouTube Shorts, transcribes them from the audio itself in more than 60 languages, and turns what worked into new ideas and scripts. Free transcript tools, no account needed: https://viewrankai.com/tools How to read any video this way: https://viewrankai.com/llms.txt