When you have an AI agent open all day, it's really easy to burn through a lot of tokens quickly. And for a while that's how everyone was using our agentic harnesses. Unfortunately, the costs add up quick.
So token maxing is dead. Whether your coding harness bills by tokens or by credits, it adds up. Let me show you a few ways to help lower you token costs! I'm using Kiro in my examples, but you can use whatever you'd like.
A lot of this builds on Laura Salinas's article Pushing Kiro's Free Tier to Its Limits. She built a project on Kiro's free tier and tracked where her credits went. It's worth a read.
I also made a video covering all five. Check it out if you'd rather watch:
(Full disclosure: I'm a Developer Advocate at AWS, and Kiro is owned by AWS. It's what I use in the video, so the examples below use it.)
1. Use a cheaper model for simple tasks
This one's a little obvious, but you shouldn't use the same frontier model for everything. Fixing a quick bug or renaming a prop shouldn't need the latest and greatest model.
In Kiro, the model picker shows a credit multiplier for each model, and everything is relative to Auto at 1.0x. The open weight models sit at the cheap end. Qwen3 Coder Next is 0.05x, MiniMax M2.1 is 0.15x, and DeepSeek 3.2 is 0.25x. Opus 5.5, a frontier model, is 2.0x, and a few models go higher than that.
The models docs spell out what that means. A task that costs 10 credits on Auto would cost 20 on Opus 5.5, 4 on Haiku 4.5, and 0.5 on Qwen3 Coder Next. That's a 40x spread for the same task.
If your model supports it, you can also turn down the reasoning effort. The docs say lower effort levels give faster, shorter responses and use fewer credits.
Heads up: Multipliers change as new models show up, so check the models page before you settle on one. Two models with the same multiplier can still use different amounts of credits on the same task, since token counts and thinking depth vary.
Keep spec-driven development in mind here. I'll come back to it in tip 3, because that's where mixing models comes in.
2. Put token rules in your steering file
Steering files are markdown files in .kiro/steering/ that Kiro loads as persistent context, so you don't have to repeat yourself every session. Here are a few token rules I have in mine. You can write them exactly like this, and Kiro will try to follow them:
---
inclusion: always
---
# Token rules
- Don't create unnecessary files.
- Don't create tests unless I ask you to.
- Try to save on token usage.
Save that as .kiro/steering/token-rules.md. inclusion: always is the default mode, so it loads into every interaction. The front matter has to be the very first thing in the file, with no blank lines above it. With these rules in place, I know I'm not getting a bunch of files or tests I never asked for.
Laura ran into exactly this in her build. With no steering file, DeepSeek generated validation files and summary markdowns she didn't need, and those cost credits. Her suggested steering rules include "Don't create summary markdown files" and "Use minimal code, no boilerplate."
Always-on steering rides along with every request too, so keep that file short. Anything framework-specific can use fileMatch, which only loads the file when you're working with matching files:
---
inclusion: fileMatch
fileMatchPattern: "components/**/*.tsx"
---
# React component rules
- Use function components and hooks.
- Type props with an interface in the same file.
There are two more modes. manual loads a file only when you reference it with #file-name in chat, and auto loads it when your request matches its description. The steering docs cover all four.
3. Use spec-driven development with two models
If you're building a feature that takes some forethought, and not just a quick bug fix, spec-driven development is the way to go. A Kiro spec moves through three phases (requirements, design, and tasks) and writes requirements.md, design.md, and tasks.md into .kiro/specs/.
The requirements and design phases are where the hard thinking happens, so that's where a stronger model earns its multiplier. Once you have a solid requirements and design document, you don't need the latest frontier model for the implementation phase. The cheaper models do pretty well working through a clear task list. You may want to try it out.
To make the switch, change the model picker before you start running tasks. The models docs say your selection applies to all later messages in the conversation. Tasks can run one at a time or all at once, so you can check the cheaper model's work on the first task before you let it run the rest.
More on the phases is in the specs docs.
4. Let ESLint, Prettier, and hooks do the deterministic work
If a job has a deterministic answer, don't pay a large language model to work it out. Formatting and lint fixes are that kind of job. Prettier gives you the same output every time, and so does eslint --fix. When you ask the agent to do it, you pay for it to read the file, reason about it, and write a diff a tool would have made instantly.
Laura's article makes the same point. She suggests hooks that auto-lint or auto-format on save, so you never spend credits asking the agent to fix style.
Start with scripts in package.json. This also includes a lint-staged block for the commit hook further down:
{
"scripts": {
"lint": "eslint .",
"lint:fix": "eslint . --fix",
"format": "prettier --write ."
},
"lint-staged": {
"*.{js,jsx,ts,tsx,vue}": ["eslint --fix", "prettier --write"]
}
}
I have a hook that runs Prettier or my Husky rule every time I'm done doing something, so I'm not wasting tokens on it. In Kiro, hooks are JSON files in .kiro/hooks/. This one lint-fixes and formats when the agent finishes its turn:
{
"version": "v1",
"hooks": [
{
"name": "Lint and format when the agent finishes",
"trigger": "Stop",
"action": {
"type": "command",
"command": "npm run lint:fix && npm run format"
}
}
]
}
Save it as .kiro/hooks/format-on-stop.json. The Stop trigger fires when the agent completes its turn, and a command action runs a shell command in your project root. The hook actions docs say shell command actions don't consume credits, while agent prompt actions do because they start a new agent loop. So for formatting, pick command.
If you'd rather run on every save, the hooks docs show a PostFileSave trigger with a matcher, something like "matcher": "\\.(ts|tsx|vue)$". File triggers only fire on changes the agent makes, not on your own edits.
For commits, Husky plus lint-staged runs the same tools on staged files.
npm install --save-dev husky lint-staged
npx husky init
echo "npx lint-staged" > .husky/pre-commit
husky init creates .husky/pre-commit and adds a prepare script to package.json. The echo line swaps the default npm test for npx lint-staged, which reads the config from the package.json above.
5. Watch your context window
The last tip is to double check your context usage. In Kiro, the context usage meter in the chat panel shows how close you are to the model's limit.
This matters because on every turn, every time you go back and forth with the model, the previous chat history gets sent again. The longer the session, the more context each message carries.
MCP servers add to this, since their tool definitions take up context too. The Kiro Powers docs put a number on it. Five connected MCP servers can mean 100+ tool definitions and 50,000+ tokens, about 40% of your context window, before your first prompt. Different harnesses may handle this differently, some use search tools to help prevent context bloat with MCP servers Kiro Powers, for example, only load their tools when your conversation mentions something relevant. Your mileage may vary, so turn off MCP servers you aren't using for the project in front of you.
Then start a new session when you move on to a new task. History from the last task doesn't help with the next one, and it gets sent along on every turn. In Claude Code, /clear does this. It starts a new session and keeps the previous one so you can resume it later. Kiro also compacts long sessions automatically, but a fresh session for unrelated work keeps things small from the start.
Wrap up
Thanks again to Laura Salinas for Pushing Kiro's Free Tier to Its Limits. It has more tips I didn't cover here, like writing a specific first prompt and pointing the agent at files with #File. The video version is up top if you want to see these in action.
Let me know in the comments if I missed one.

Top comments (3)
Are you wasting tokens?
erik, this is a highly practical breakdown. tip #4 (letting deterministic tools like prettier and eslint do their job) is the biggest token-saver that developers constantly overlook.
as someone building entirely on a $150 phone, i rely heavily on strict free-tier api limits (like groq or openrouter). your point about using cheaper, open-weight models for simple tasks (tip #1) is exactly how i survive. i reserve the frontier models only for complex architecture planning, and use lightweight models for boilerplate or simple refactoring.
one addition from the mobile constraint regarding tip #5: clearing the context window isn't just about saving tokens. on a budget android browser, rendering a massive chat history with 50k+ tokens of context will literally lag or crash the tab. starting a fresh session is a hardware necessity, not just a cost-saving measure.
thanks for sharing these actionable, no-fluff strategies! đŻ
Really liked this because AI coding costs are easy to ignore when youâre focused on getting things done. The âuse the right tool for the jobâ idea is probably the biggest takeaway for me.
I especially liked the point about not asking an LLM to do deterministic work that ESLint, Prettier, or a simple script can handle instantly. It sounds obvious, but itâs surprisingly easy to fall into the habit of letting the agent do everything.
The context-window point is also underrated. A long-running AI session can feel convenient, but carrying all that old context into every new task isnât necessarily helping. Sometimes starting fresh is both cheaper and actually clearer.
Overall, I like that this isnât really about using AI lessâitâs about using it more intentionally. Spend the expensive reasoning where it actually adds value, and let traditional tools handle the boring, predictable stuff. Great practical breakdown.