# GLM 5.2: What you need to know Channel: Greg Isenberg Video: https://www.youtube.com/watch?v=xa-9O5cDm3c Duration: 23 min Language: English Words: 4437 Transcript page: https://viewrankai.com/tools/youtube-transcript/xa-9O5cDm3c --- [0:00] Okay, you've probably heard of GLM 5.2 that's going viral everywhere on Twitter. Yes, it's this new open-source local AI model that people are saying is the chat GPT moment for local AI. But, no one's actually gone and shown you how to use it and how do you actually set it [music] up? So, I figured I'd bring on my friend Amir. He tells us exactly how you should think about [music] running GLM 5.2, how you should think about running local models, how that integrates [music] to something called open router, how you can use it with your codex or cursor or cloud code. In this episode in 20 minutes or less, [0:37] you're going to get everything you need to know about local AI models, why GLM 5.2 is crushing benchmarks, and how you can set it up today so you can go and build your startup, build your business, [music] and be more productive, and be more efficient. Enjoy the episode and I'll see you at the end of it. Hit a like, comment, and subscribe for more of this [0:59] sort of stuff in your feed. Enjoy. [music] [1:06] [music] Welcome to the show, Amir. Uh by the end of this episode, what are we going to learn? [clears throat] We're going to talk about especially learn about how local models are kind of keeping up now with the pace of these closed models as well and how you can kind of use compounding models or fusion models as as open router calls it to be able to do sequencing between a more extensive thinking model and a more execution-based model. We'll show you how GLM 5.2 actually compares and stacks up against other models and how you can effectively use it and get set up with [1:39] it as well. Cool, I did an episode on local models. It was a hit. People wanted more tactical, how do I actually, you know, how should I think about local models? How do I implement local models? How I can actually make this a part of my daily workflow? So, I brought on Amir. Welcome to the show and let's [1:59] I give you two welcomes, by the way. That's how excited I am to have you share share with everyone everything. So, let's get into it. I'm super excited to be here. Let's jump right into it. So, let's talk about what happened this week. Google came out with Gemini 5.2 and I think this was a big inflection point because we've typically seen with local models, it's either, you know, storage extensive that you can't essentially run it or install on your computer or you need a better like GPU RAM performance to be able to actually [2:28] run it locally as well. Now, Gemini 5.2 is also resource extensive, but what we're seeing is with open-source models like open-source providers like OpenRouter or Llama being able to help you run these models in the cloud and effectively being able to essentially pay slightly less for input and output tokens compared to the more closed models. Now, what I want people to take away from this session is one, how to actually get set up with it. We're not going to go through the detailed setup process, but I'm going to just cover how you can do it in Cursor using OpenRouter or in Code X and then effectively talk [3:04] about how Gemini 5.2 stacks up against the other models and then how you can effectively use it as well to do model chaining. And then we'll do model chaining and then we'll do just a kind of a maybe a quick walk-through of how I'm currently using it. You know, I want to be very honest like these local models [clears throat] still have a lot of work to do in terms of having tool capabilities to be able to, you know, have the model, you know, the modalities to be able to see images and conceptualize on what they're looking at and I'm going to tell you how you can effectively [3:34] circumvent some of that where you can use other models to explain what the the image is back to Gemini and then have Gemini work on it. And then also just have a very live test on how this stacks up against other models. You know, benchmarks are great. Personally, I'm not an expert in it. I don't know what any of these benchmarks actually mean. The way I do it is off of like let's build it out and see how this actually looks and how it stacks up [3:57] against other models. Sound good? Yes, sir. Let's do it. Okay, so GLM 5.2 came out and essentially has a 1 million context window and it scores 81 points on the terminal bench 2.1. It's just about four points behind Opus 4.8 and it does quite well on the long horizon task evaluation. So this is essentially projects you have that you want to run long sequence tasks on and you know, I think it's seeing account like the thinking parameters and how it can think through and plan through some of the the task at hand. So you can see that across all these different kind of benchmark [4:37] reporting reports. GLM GLM 5.2 actually does quite well. So in this case it's 62.1% compared to Opus's 69.2. What's special about GLM is it's open source. You can run it locally on your machine if your device can support it or you can run it in the cloud through the open model providers. It's a big leap from 5.1. I personally didn't test 5.1. I got straight into 5.2, but from what I'm seeing based off of like Twitter and conversations with people, it's performing quite well especially on the front end side of just execution based tasks. I haven't really tested more on the back end resource intensive tasks, [5:16] but just based on perception and what we're seeing in the reporting, it's stacking quite well. So when you say stacking well, we talking like cuz like I look at benchmarks and honestly it goes through you know, one ear one eye and it goes out the other eye. I'm like I glaze over cuz I'm like what does this really mean? Like are we talking is it like 4.8? Is [5:37] it like 5.5? Yeah. What do we How should I think about this and yeah, give it give it to me straight. Honestly, man, I'm in the same. boat. don't get it. You know what I mean? Like I'm I'm not smart enough to understand how these benchmarks actually stack. So, I'll be honest, for me it's like let's just let's just build it, use it, and see how we feel about, you know, it how it performs compared to the other models. And for me, I want to get the best out of it. So, if I feel like GLM 5.2 is strong in one part, but weak on the other, then I think about how do I [6:08] actually use other tools or other models to essentially Now, I think like almost like a fusion approach and I think Open Router which is one of the like the model providers coined this where it's like you're able to do like sequencing between two different models to get the best output. So, I'm totally game on. If I can run a local model on my machine to do certain tasks, but then call, you know, Opus or Codex to do something else and have them work together, by all means. I want to be the most token and [6:38] cost-efficient and performance as well. Yeah. So, on the setup side, I personally started using this through Cursor using Open Router's API. So, how it works essentially is you got to go to Z AI, which is the GLM provider. They created the GLM 5.2 model. You get an API key from them, and then you take that key and you go into your Cursor settings, paste it into the OpenAI field, and then from there you override the OpenAI endpoint with this API endpoint right here. And essentially from there, you go back to models, add a custom model, GLM 5.2, and you're able to then actually call GLM 5.2 directly. So, in [7:23] essence, instead of OpenAI key, you put your GLM API key from Z AI, then you override the API endpoint for when you call OpenAI chat completion with this one right here, and then you go back into custom models and add the custom model protocol. You You also alternatively do this using Open router. So, if you want if you're using codex, you can go to open router, get your open router key, and then um go into the provider, get the endpoint, and then go into codex, create a profile, and say, "Hey, I want you to install this model um open source model." Codex does actually support open [7:59] source models. So, you're able to provide the details of what the model is, the context window, and then essentially when you're running codex through the CLI, you can switch to GLM 5.2. Easy enough. Yeah, easy enough, you know, and then maybe we can uh have a page or something to show later on where they can kind of follow these instructions. There's a lot of guides on Twitter and online. You can follow those, but essentially in my opinion, the best way to get started with this is just go to open router and cursor and get that set up [8:28] through and through. So, um let's talk about here the model. We've talked about the benchmarks, how it performs. Really, if we want to just at least take you know, put some weight to the benchmarks, I'd say if they're scoring at 62% and Opus 4.8 is 69, you know, that probably means something, you know, for for for the normal people, we probably won't really know until we actually play around with it, but I went in and was just looking at, for example, this website we have, this is a small app, and I built this in I think Opus 4.8. I was just testing it around, and I started refining the [9:02] design using GLM 5.2. So, I was like, "Hey, redesign the hero section for me or refine it." Um there's this like section right here with all these images. I was like, "Why don't we just do a little like carousel style?" So, it you know, it's fascinating in a way because I don't I personally don't think open uh the the local models had um this kind of capability to be able to get get it so refined and accurate um [9:29] previously. And I I I tested the models like that we had, um and I find that GLM 5.2 is a lot more refined, and it's able to follow the instructions on like what you want it to do. So, in a couple prompts I was like, "Hey, you know, let's do a carousel here. Let's make sure we, you know, are able to show the the images and then from there I want you to build out like a Bento grid style of all the features that we have." Now, this is all one prompt. Obviously, you know, you can see it's a little bit of vibe coded here. It has a little side like [9:59] badge the labels, you know, and and you can tell, but at the end of the day for a local model for like, you know, a local model if you're running this on a computer and not burning any tokens, it's still it's doing quite well and I think um I can see how it stacks like in terms of [10:16] reporting. Yeah. I mean, I think so like, okay, what is the what is the the main benefit of using a local model versus something in the cloud is you don't burn tokens. You, you know, you essentially the way to think about it is you're buying a machine um and and correct me if I'm wrong, but you're buying a machine um we can we should talk about like some of the machines that people could [10:42] potentially buy. But you're buying a machine, it's a you know, one size it's a one it's a cost. It could be 2,000, 5,000, 10,000 and then you can just run tasks, right? So, you're building a startup, you just want maybe you want conversion rate optimization. So, maybe you say every day I'm going to feed you customer feedback and every single day I want you to work on the front end and you just do that. So, my question to you is for local models specifically, for people who actually want to be building companies, shouldn't they be basically running it all the time on certain task and how should people be thinking about [11:20] it? Yeah, so I think this is that's a little bit tough to answer and I'll say why because like this model specifically is really resource intensive. Intensive. So, a lot of I think um existing consumer uh computers may not be able to run this from what I've seen. I've been I've been running it on the [11:38] cloud through open router directly. Um, And what's the cost with that? Yeah, so uh with So, I actually was trying to map out the token cost of model training. So, if we had about 50,000 input tokens and 85,000 output tokens, uh to get almost close to an Opus 4.8 level of output, it will cost us 44 cents. Whereas with Opus 4.8, it costs you [12:05] $2.38. So, the you know, there's a big almost like a big difference on the almost like 5x, you know, price difference between Which doesn't sound like a lot when you're like, "Oh, $2 here, 44 cents here." But, when you're actually using these these things and you're running it all of the time and you have pretty, you know, big tasks that you're going after and you don't want to be constrained by [12:30] token costs. So, it's a big deal. 5x is a big deal. Yeah, yeah. And this is based on kind of the averages on like the coding benchmarks and what what cursor charges you through the API pool. But, I want to I want to I want to be also future-proofing yourself. If we got this far with Gemini 5.2, I wonder what next six months looks like, right? So, would it be even worth like I think it's also worth thinking about making the upfront investment in your machine right now to be able to potentially download and run these local models so that when we get to Gemini Gemini 5.3 or 5.5, [13:05] we're essentially made the upfront investment in the in the in the compute and the equipment to now save a lot more in the long run with other future models that may come out that are going to be much more expensive. Cuz I think this, you know, we're seeing the the the AI subsidy on tokens, right? Like we're getting a lot more output out of Claude, out of Codex. And I wonder, you know, I I've personally seen it, I'm sure you have too, where it's like now we're hit hitting our usage a lot faster than we before especially when Fable came out I ran I remember like I ran it in like in [13:33] the first day I hit my limit you know. Totally. So you're what you're saying is basically like it you know if you look at uh I mean if you look at the history of VC back startups think about Uber when Uber first came out they actually subsidized rides and they got you hooked onto the onto the onto the app and then over time they started increasing prices increasing prices. What you're saying is with uh you know in the AI age with a lot of these LLMs they're going to get you hooked into the workflows you're going to you're going to build on top of it and over time [14:07] you know those subsidies are going to go away as they go public and things like that. So what you're saying is maybe it's a good idea to actually invest in running this thing locally um because you know the price of memory isn't getting cheaper. Mhm. And the price of tokens aren't getting cheaper. Mhm. So building it now securing it while you [14:32] can might be a good idea. Yeah and to tie this all together I'd say two things. One um harnesses that are agnostic on models so like for example cursor where you're able to run multiple models across the same sequence of tasks are going to actually potentially benefit from this right? So you know I wouldn't be surprised if one you know cursor decides to directly support Gelion 5.2 as a model provider and lets you kind of tap into that cost saving um if you couple it with like composer 2.5. So this is [15:04] where it goes into model training right? Where I was for example looking at um earlier in this task here I wanted to refine the hero section. So what I did is I actually used Opus 4.8 to um to first import screenshots because Gelion 5.2 doesn't support vision capabilities. So, what I did is I actually used Open 4.8 to import screenshots and explain back to me what it sees, right? I was like, "Tell me what you see specifically on the front-end design for the hero section and lay it out." And then I switched to GLM 5.2 to study that layout and then actually act on making those [15:44] changes. So, it's kind of a way to like circumvent the fact that you have limitations to what GLM 5.2 can do 5.2 can do like image capabilities, but you're able to kind of now chain the expensive model to think through the plan and then get the same level of frontier like frontier level like quality, but at a at a much [16:07] affordable price point. I mean, that makes sense. So, your your recommendation is basically you know, there's it's almost like free trade versus protectionism, you know? The world you you you know, not to be this isn't political. This is just economics theory, right? Which is like, you know, when when when people are trading with amongst each other, you know, maybe it's you know, in Canada where you are, you know, you might want to trade with Florida cuz you know, we got good oranges here. And we might want to trade, you know, we we don't get we can't make maple syrup here. So, we'll [16:44] get your maple syrup, right? your maple maple syrup. Yeah, exactly. So, make the best use of it, yeah. Make the best use of it. So, you're saying is, you know, using cursor and and basically you know, you don't have to use cursor, right? You can use whatever you'd like. Codex, Claude code, yeah. Use one of those to basically say like, okay, for certain task I'm going to be using local models, for certain task I'm going to be using the best-in-class cloud models. And then together, ultimately you're getting, you know, great results in terms of the output, but you're also not spending through the wazoo. And if you're if you're a token [17:19] maxi or like you and me are, like in the sense of like we're we're always pushing into the limit around anything we're building to get the most out of AI cuz you know, we don't want to hire 100 people, 500 people, and stuff like that. Um it's helpful to do that. And exactly and anecdotally to two parts, right? One internally within our company, you know, I think Satya at Microsoft, you know, mentioned how like human capital plus token usage is now a big factor into what they're doing, right? A lot of companies are now moving away from, you know, having direct access to the cloud code API to run the [17:56] tokens cuz of how expensive it's become, right? So they're canceling subscriptions. So we're seeing this first hand with a lot of companies now are saying, "Okay, cool. This first year was great. You know, we had the mandate, you know, AI adoption, token maxing. You know, that's how we're going to measure success and that's how we're going to become AI native." Now they're like, "Wait a minute. Okay, cool. We've done We've done this, but we're spending way too much money on tokens. How can we now be more effective, right?" And I'm seeing this first hand, too, where it's [18:21] like, especially now, right? In a way, you can have some sort of direct ROI between the tokens you're spending within the engineering team cuz you're like, "Okay, cool. We're saving a lot of time. There's an output. You know, engineers are expensive. We get that. But now you're providing the same level of harnesses and models to the non-engineering teams that, you know, are one-shotting a like, "Hey, help me format this email." and they're using Office 4.8 AI thinking. They're like, "Maybe that's probably not the right model." and that's a governance issue, right? That's a big You know, I'm having these conversations with companies right now where they're [18:53] saying, "Hey, can you help us figure out like how to build governance and proper education on how to actually use the right models?" And this is where I think model chaining is a big factor, right? By the way, you know, John at marketing, maybe you shouldn't use Office 4.8 to run this like to just format this email for you. And just helping them understand that. And I think, you know, I wouldn't be surprised if in a year from now companies start You know, we've been thinking about it as well. We're like, "Hey, why don't we just get our own machines and start running some local models cuz it's a lot [19:20] more effective, especially how much money we're spending on tokens." What's the like just to play devil's advocate, why wouldn't I just use open router and call it a day? Please do. Yeah. Yeah, absolutely. I think they should. cuz I you know, when I'm on on X, I see a lot of people being like, "Buy a Mac Studio or buy, you know, these expensive [19:40] devices." You know, for for people listening, do they need to buy a local piece of hardware? Like if the price even goes up 2X or should they just use open router and cursor or or or Claude code, you know, whatever harness they they want? If I just have this right machine, I'll get to this you know, result. That's not how it works. You know, no, you don't need a Mac Mini. You don't need this [20:05] equipment. You can get started today. You know, and what I love about open router and all these other tools is again, they're so agnostic. They make it easy for you to be able to access this in the cloud. They run the models locally. And you it's credit-based. Load $20, get it going, and it's easy to set up. Um I highly highly recommend if you're starting to dabble with this and token usage is a is a thing for you, get one of these Asian harnesses set up now that they're a lot of them are model agnostic. Run some tokens in open router, get these open models in there, [20:34] and start start vibing. Just get some like I love to experiment to see how far I can take this. What if I fine-tune with review with Gemini 5.0 execute 5.2 and then review with composer 2.5 or codex 5.5. There's a lot of ways and I think we can be really effective and I think that's what the smart people are going to be doing in [20:53] the near future. Uh some people are saying, "I don't care how much tokens cost because I think there's so much opportunity in building startups and and optimizing and AI arbitrage that I don't even care if it costs me whatever." What do you say to those people who are just basically ignoring this whole open source local AI [21:14] movement? I used to be the exact same person, you know, in our first episode you're like, "Yeah, how much does this all cost?" I'm like, "I don't know, man. I'm just vibe spending." And I think my that mentality has changed. Now I can see that my usage limits are being hit faster and my cost is going up and now that our team is expanding internally as well. So, I think as a solo person it's a lot easier to build a a case or rebuttal on why you should just token max as much as possible, which itself is kind of a like an it's ironic. You shouldn't be token [21:43] maxing. You should be token minimizing as much as possible and output maxing instead. So, my answer to that is if it works for you and you can directly have an ROI that you can show that, "Hey, I spent $200 and got a thousand out." Great. Otherwise, sooner or later the subsidy is going to run out. All right. Well, I think that's that's the episode, you know, unless there's anything else you want to add before we [22:05] bounce. Yeah, I mean, for people that are trying to dabble with it, play around, have some, you know, see see what it can do at least in the front end for you and start working back end tasks and yeah, I hope they they got they learned something from this. I'll include links for where to follow Amir. Um he's always one of my first calls whenever I'm trying out new stuff and so I'm happy that you were able to jump on. I appreciate you. We appreciate you. Give give him a follow, like, and comment [22:30] this video. Let us know what you think. We'll be in the in the comment section. Uh just just you know, out there trying to help and learn and and uh and uh thanks a lot, Amir. I'll catch you on the next one. Thanks for having me. --- About this transcript Read from YouTube's own caption track and laid out by ViewRank AI (https://viewrankai.com). ViewRank AI finds the videos already beating a creator's own average on Instagram, TikTok and YouTube Shorts, transcribes them from the audio itself in more than 60 languages, and turns what worked into new ideas and scripts. Free transcript tools, no account needed: https://viewrankai.com/tools How to read any video this way: https://viewrankai.com/llms.txt