# Paperclip AI Cost Management: How to Stop Burning Tokens (Real Settings Walkthrough) Channel: Fru Dev Video: https://www.youtube.com/watch?v=UIdH5Ac1Db8 Duration: 19 min Language: English Words: 3411 Transcript page: https://viewrankai.com/tools/youtube-transcript/UIdH5Ac1Db8 --- [0:08] Greetings everyone. Fruit here. I have 15 agents working 24/7 to optimize my life and I make videos to show you how. Now, Paperclip is a tool I talk quite a bit about and if you've watched the channel, you would understand how to get Paperclip set up and to have your own army of agents working for you 24/7. Now, when you use Paperclip, cost is a topic that comes up very frequently and it's one of those things that Paperclip is a fun tool to work with, but when the cost conversation comes up, it stops being fun and it could be very painful very quickly. So, I want to [0:42] spend some time talking about some cost considerations to be aware of when using Paperclip and some settings you want to pay attention to from a cost perspective. Now, coming in here in Paperclip, if you're not familiar with this, so many videos on the channel, check them out. For cost in particular, it really comes [1:00] down to the agent you're working with. So, if I have an agent, every time that agent runs, it does use tokens. It does use tokens and you have to provide it a model to use. So, if I come into my configuration, picking this particular agent, which is a business agent, what I'm going to do is scroll down. You can see this adapter in here. So, this adapter tells the agent what model to use. In here, I am using the Claude local. What does that mean when I say using the Claude local? Now, I'm going to go back over to the CLI. If you have Claude, you probably know what the Claude CLI [1:36] is. So, in here, I'm in my Claude CLI, I'm authenticated. I have the max plan for Claude. Because I'm authenticated into this with the max plan, it uses that. Actually, it says API usage billing. Now, I have the max plan and I switch between accounts. So, it it would use it would use that billing. If you have the max subscription plan and you come in here, if you authenticated in there, you select the local adapter, it would simply pass [2:06] through and use your subscription plan. Now, if it's using your subscription plan, one thing that folks don't pay attention to is they're quickly run out of out of their credit limits for the week or for the day so fast that it becomes concerning. And here's a couple of things you need to do to address that. Number one is really thinking about your heartbeat. So, if [2:28] you have an agent, so I have 15 agents. Imagine if all these agents are running every 2 minutes or every 5 minutes. That is a lot. 15 * 5 it's it's it adds up pretty quick. Now, imagine if those agents are all running every 5 minutes using and I'll go back up here, using Claude and the model of choice is, you guessed it, they are using Opus 4.6. So, 15 agents running every 5 minutes using Opus 4.6, you're going to get yourself into trouble pretty quickly by burning all your tokens. So, that could be that is [3:02] definitely not something you want to do. So, how do you solve that? Number one is think very clearly about your heartbeat. So, heartbeat is how often this agent runs. So, if you don't want it running every day, don't put this to 5 seconds. Because if you put five here, that's 5 seconds. If you put 60 here, that's [3:19] that's running every minute. So, if you don't want that, if you don't want it run every minute, make sure you put a big enough number here. So, if you put 360 here, that is every hour three This is zero in front can't seem to get out. So, this is making this agent run every hour. An agent that runs every hour will cost you less than an agent that runs every 5 minutes. I think [3:43] that's pretty self-evident. Alternatively, an agent that doesn't run at all, it will cost you way less than an agent that's running every hour. So, your your cadence becomes important. For some agents, you might just want to keep them not to run and to only run on a heartbeat or manual run. So, you're coming here, you you hit manual run. That's the only time you want your agent. So, if I think about estate, you maybe you do your estate planning once a year or once every 6 months. Is there a reason for this agent to run every day [4:13] or every week? Probably not. At that point, you want to figure out what what's the cadence for every 6 months and then Um maybe for your chief of staff that needs to run maybe daily and send you your daily brief, then sure, set it up to run every day. But making it run every hour, is that really necessary? Unless you have a I used to have an agent called content agent and that agent would [4:35] I think I called it the scanner agent. It would run every hour in my domain, in my space. As you can imagine, my my job here is to stay on top of AI and and news and things that are happening. It would essentially scan the web for news and certain file um uh certain profiles and send me a brief [4:52] every hour. But that that was very costly, too. So, I kind of stopped that and I went to a morning brief and an evening brief and that helped to manage the cost. So, two things. Number one, pick the right model. And then number two is set the right schedule. That will help significantly with [5:12] uh cost. And then number three, too, when the agents are running, you want to think about your skills. And this was wasn't very important very clear to me uh which is I tend to write I I use the agents to write my skills files and if you're not sure with what a skill file is or what the agent instruction file is, again, I've made videos around this. This is [5:34] just the boss. It's just a lot of text. I mean, you can tell that AI wrote all of this because it did. AI did actually write all of it. But the reason why I'm saying that in tongue in cheek here or maybe not tongue in cheek is because all of this is tokens. Every time this agent runs, it packages all of these instructions as tokens and sends it to the model. And you're paying for those tokens in and you're paying for those tokens out. So, is there a way to still convey your instructions without having this just burned 5,000 tokens right off the bat, right? How many [6:04] However, how many words this is. So, thinking about your your agent instructions or your skill instructions and how many skills you give to an agent. Are you going to give the agent all skills even if the agent doesn't use all the skills? Those are things you have to think think about. So, in this case, most of my agents, I just need them to update my calendar and maybe send me an email. That's it. Give them those skills because this skill instructions, at least the definition of the skills, goes into the agent and that still consumes some tokens. All be [6:35] it. So, something to be to be aware of. Now, I think it's Peter Drucker or somebody that said to the effect Peter Drucker is a good was a manager like he was a very good author, wrote a lot about management, said you cannot manage what you don't measure. So, the reason why I'm saying that is if you come in here and I've talked about cost, you might say but through, well, that's interesting. I see you run some models here but what does it what did it cost you? So, pick on this agent courier agent and um this agent is as run a couple of times in this demo [7:17] environment and you can see the input tokens here not as much. Output tokens quite a bit. Cash tokens 6.2 million. That is a lot. But the thing is cash tokens is is actually cheaper. It's about 10 times cheaper for cash tokens. And and I can talk about cash tokens in another concept in or in another video. But why is the cost zero? Are you [7:39] telling me that this is costing zero? No, it's not. The reason why it seems to cost zero is, and I'm sure Paperclip will fix this, is if you use a subscription and I run all of these models on the subscription on my Mac subscription. For some reason, it doesn't and I don't know what the explanation is, but I think it doesn't calculate the actual [7:58] cost. But, if if I was logged in here with API, it will if and I run I didn't have it run this today, but if I was running here with API, it will um it will actually show you the cost. And then, at that time, one thing too I should I should go back is let's go back to the agents. Let's go up to the agents. Um you want to think about the budget. So, most of your agents, just talking about cost, set a budget. Whenever the budget hits, it [8:26] will uh stop this agent from working. But, the budget will not hit if we are dealing with a situation where uh the cost is zero, right? How is the budget ever going to hit $5 if the cost is always zero? So, I found a workaround. I'm not sure my workaround, but I'm really calling it a workaround because I'll be surprised if the Paperclip folks don't fix this pretty soon. Most people who use this don't use the API. They use the subscription. And my big reason for using Paperclip is because I could use my subscription [8:54] tokens. Not so with not so much with Openclaw, at least as far as I know. So, I'll go back over here. I'll tell Paperclip, "All right, can you help me query Paperclip's embedded PostgreSQL SQL directly and uh you can use port 54329 particular and uh understand the usage data on what is costing me uh to run Paperclip, even if I'm using the the uh Anthropic subscription as opposed to the API base uh from from Anthropic and then apply Anthropic's API pricing to the token counts to show me what is costing me to run Paperclip using Anthropic's uh subscription. And then always backfill the cost uh in cents on paper clip so I [9:43] can see on the UI in paper clip how much each agent is is costing me. Uh if you need to write code to achieve that uh do it, but I want to see the results uh right now. So, that's the instruction I'm giving. You can maybe pause uh pause the video, take a screenshot of it, and I'm hoping that this instruction works because I tried it on my other environment and it seems to work to show me the cost. To a certain extent, I'm hoping that with the prompt in here I will um we will be able to um to get it. My the way it worked when I [10:17] did it was it actually wrote a script that will query because paper clip behind the scenes is using Postgres for actually running the application. It wrote a a Python script that will query that instance, uh figure out the tokens because I mean it's you when you make those calls all of that is stored in the database, and applied the the then it applied. So, it So, it's seeing the cost events database is looking at that. So, I think this will work. While this is working, I want to go in and touch a little bit about this tokens in. So, you can see tokens in is pretty small. Uh and the [10:53] reason why this is small here is it simply for this particular agent it simply wakes up and tell the agent to run. And let me pick up a particular run. Let me go up here. I think this might be a good one to talk about this particular topic. So, if I go in here, uh you look at this particular run, you can see that when it the token in is so small, this is 5.4 5.4 54 tokens in. That is tiny. Tokens out is 4.8 thousand. It's because token in is essentially just triggering to say wake up agent run, right? It wakes up this [11:35] current agent and says go ahead and run. So, the agent when it's running it it kind of is doing a lot of things. It's uh you can see it's setting up the context, pointing to a lot of uh stuff, setting up the environment. And then uh it goes into this really verbose output that that comes back authentication. I think it ran into some authentication issues in there. So, all of this I won't uh belabor the point here with showcasing all of that. But, all of that uh contributes to this output tokens. Remember, you're paying for tokens in and you're paying for tokens out as far as these models are [12:09] concerned. You can look at the pricing. Now, what is a cache token? The cache token is as you as there are several runs um there is a folder and I wouldn't show it right here. Claude has this concept of the cache tokens where it saves uh the the the things that it's getting back. So, it's not always hitting and that is that's my understanding of it. It's not always hitting the Claude API every single time. And that is considered cache tokens and those cache tokens builds up with more runs you have. So, but you don't Now, I was very terrified when I saw this to say wow, means I'm paying [12:46] for this amount every single time. No, you're not. You kind of paying for it, but you it's not at it's not at the same cost as this. It's about 10 times At least that's what I heard or read. That's about 10 times cheaper than the most expensive is this and then this is not as expensive and then this is way way cheap. So, but then it helps with the context and managing the context with longer conversations. And then over [13:10] time Claude clears up the cache tokens. You can you can go delete it yourself in the dot Claude like there is a folder in there, but that's a little bit of an advanced concept I don't want to be getting to right now. All right. So, let's see it's um it's run. And I told it to do it right now, so that's the cost. Let's go back [13:31] to UI with any log. So, if you don't see the cost, I'll refresh. Like I said, it just took a Nope, it doesn't show. It doesn't show. Does the Chief of Staff show anything? All right, I'll say All right, I see you show me the numbers here, but why isn't it showing on the actual paperclip UI? I've refreshed the [13:55] UI, but I don't see anything. All right, so I'm asking. So, in here it's telling me the agent. So, Chief of Staff is $2 based on the inferred cost. Um Carrier is 2.4, calendar agent is 1.9, and then I have some agents while the Omega Dev. So, I have a couple of other companies in [14:16] there for demos. Um so, those are kind of costly. But, why? I'm expecting now cuz what I'm expecting is to see the cost showing up in here. Um so, that should take a second. Yeah, I really wish I didn't have to do this uh to get the cost. Okay. What else uh to touch on from a [14:39] cost perspective? Oh, the the last piece and I really had this in my notes. Uh I would have been uh I'm disappointed if I didn't touch on this is when you set up your agents from a cost perspective because this video is about cost. Not every agent needs to run on Claude Sonnet 4.6 or Opus 4.6. Not every agent needs to run on that. As a matter of fact, not every agent needs to run with Claude. So, this is where you can have and I actually have a setup where I have um or some local or Llama models running and I could give them like a very basic [15:21] I I don't have my Llama on this particular machine. I can give them like a basic model like QN, you know, 8 billion parameter for like the web search model. And then because in that case you're not paying and I have that running every, you know, hour. So I'm not worried about paying for that. And once it does its web search, it's writing that into a file and then once a day my chief of staff when it runs, it's [15:47] it reads that file. And so it's not spending time with the web search and all of that. It's just summarizing that and making it more intelligent for me to consume and integrating that into my entire uh system. So that is one way. So from a cost perspective, I think we've talked about your the model you choose, right? The type of model you choose, the cadence of running the models, managing your tokens or or managing your your your context uh to tokens in, tokens out. Uh that goes into how you structure your instructions. Uh those factors all all come in. Putting budgets uh so that uh when things go awry, you don't run [16:29] into issues and I think we finally seem to be having uh some numbers show up here, which is good. Uh let's see. Shows up here. Let's go back. Does it show up on the dashboard? Not yet. So um but if if you don't see numbers in here and you're wondering why why don't you see numbers, like I said, it's because if you run with the subscription, [16:54] then the numbers don't show. Uh honestly, if this doesn't refresh at this point, I would just say keep prompting. It should work. I got it working in another in another session. I was really hoping to see the total cost here. But somehow it's not showing. Uh but I see it in here. So it's showing [17:15] it here. Observe budget. Um So, I would just keep prompting and this is what I I tend to use paper clip with cloth quite a bit as you can tell. I was prompting here to say, "Hey, I'm not seeing it. Can you go fix it?" And eventually it will fix it, but I thought I would show that to to you guys. So, um hopefully cost is not your concern. I think somebody in the comment section below today uh asked that question on they used paper clip and it just blew away their weekly limits. Well, [17:44] hopefully this gives you some levers. There are many levers you can use on this. Hopefully this gives you some options to uh consider and uh use to help with that. And um yeah, as always I'm definitely reading the comments. If there's any topic I should make um make some videos about, let me know and I'll see what I can do. If you haven't gotten the vault, you can grab the vault. It's a starter pack. Use it as a startup. There are about 200 prompts in there to give you some ideas on on how to ask questions against your data. Um and even if it's not relevant to you at this point, look at [18:17] it as a way to support the channel as we're working and growing and building AI to help you be AI ready. AI ready you. Thank you for watching. As always, this is Fru. I'll see you in the next demo. --- About this transcript Read from YouTube's own caption track and laid out by ViewRank AI (https://viewrankai.com). ViewRank AI finds the videos already beating a creator's own average on Instagram, TikTok and YouTube Shorts, transcribes them from the audio itself in more than 60 languages, and turns what worked into new ideas and scripts. Free transcript tools, no account needed: https://viewrankai.com/tools How to read any video this way: https://viewrankai.com/llms.txt