AI Cost Control in Production: What I Learned Building CVChatly

Photo by Homa Appliances on Unsplash
The bill that changes your architecture decisions
When I shipped the first working version of CVChatly's AI pipeline, I was genuinely excited. The agents were orchestrating well, the LLM responses were good, and users were getting real value from the job search features. Then I looked at the token consumption for a single user session and did the math at scale. That number made me sit down and rethink several decisions I had considered already closed.
This is the moment most teams hit somewhere between prototype and production. The demo works. The unit economics don't. And unlike a server bill that grows linearly with users, LLM costs can compound in ways that are genuinely surprising if you haven't thought through your agent design carefully from the start.
I'm not going to give you a generic list of tips. I'm going to tell you what I actually changed in CVChatly's architecture and why, because I think the reasoning matters more than the checklist.
Where the money actually goes in an LLM agent system
Before you can control costs, you need to know what's driving them. In a multi-agent system, the token spend comes from a few distinct places, and they behave differently.
- System prompts repeated on every call. If you have a verbose system prompt and you're calling the model frequently, that context cost adds up fast. I had one agent with a system prompt that was nearly 800 tokens. It ran multiple times per user session. That's a lot of context you're paying for on every single invocation.
- Conversation history passed as context. Agents that maintain conversational memory by passing the full history back to the model each turn can balloon in cost as sessions get longer. The tenth turn of a conversation is paying for the first nine.
- Redundant calls for tasks that don't need LLM reasoning. Early in the build, I had the LLM doing light formatting and classification tasks that could be handled deterministically. That's not a product decision, that's just waste.
- Model selection by default rather than by task. Using the most capable model for every step feels safe. It's also expensive. Not every step in an agent pipeline requires the same reasoning power.
Once I mapped the call graph and tagged each node with its token cost profile, the picture became much clearer. Most of the spend was concentrated in two or three places, not spread evenly across the system.
The architectural changes I made, and what I'd do earlier next time
The first thing I changed was system prompt design. I went through every agent prompt and stripped everything that was implicit or redundant. Good prompt engineering isn't just about getting better outputs — it's also about being precise enough that you don't need 600 words to say what 150 words can say. I also moved shared context that multiple agents needed into a structured object passed at the orchestration layer, rather than embedding it verbatim in each prompt.
The second change was smarter context windowing for conversational agents. Instead of passing the full message history, I implemented a summarisation step that condenses earlier turns into a compact state representation once the conversation passes a certain length. This does introduce some complexity, but the cost reduction justified it quickly.
The third change was model routing. I mapped each agent by task type and reasoning demand, then assigned models accordingly. The agents doing complex multi-step reasoning and judgment calls use a more capable model. The agents doing classification, extraction, or structured formatting use a smaller, faster, cheaper model. This required some testing to validate quality didn't drop, but in practice the lighter models handled their scoped tasks well.
The fourth change was caching. For certain types of queries where the input is predictable and the output doesn't need to be freshly generated each time, I introduced a caching layer. This is particularly effective for things like job category classification or standard document parsing where the same inputs appear frequently.
What I'd do differently from the start: instrument everything before you build. I added proper cost tracking and per-call logging later than I should have. If you're building an LLM-powered product, treat token spend as a first-class metric from day one, the same way you'd track API latency or error rates. Build your observability layer early, not as an afterthought when the bill arrives.
Cost control and product quality are not opposites
There's a temptation to frame this as a trade-off: you either have a great AI experience or you have manageable costs. In my experience, that framing is wrong, and it's actually a sign that the architecture needs more thought rather than a fundamental tension you have to accept.
When I tightened the system prompts, the agent outputs got more consistent, not worse. When I introduced model routing, the faster models on scoped tasks responded quicker, which improved the user experience. When I added caching, repeat interactions felt snappier. Cost control forced me to be more precise about what each part of the system was actually supposed to do, and precision in system design tends to produce better products.
The one place where there is a genuine trade-off is in how much reasoning depth you allow on complex tasks. A more capable model with more context will sometimes produce a meaningfully better output. The answer isn't to always choose cheap — it's to be deliberate about where the quality delta matters to your users and where it doesn't.
In CVChatly, the moments where a user is getting personalised guidance on their job search positioning are worth the cost of a more capable model. The moment where the system is deciding which of five predefined categories a job posting belongs to is not.
Practical starting points if you're in this situation now
If you're building or running an LLM-powered product and haven't done a cost audit yet, here's where I'd start:
- Log every LLM call with its token counts, model used, and the agent or feature that triggered it. You can't optimise what you can't see.
- Identify your top three cost drivers. In most systems, a small number of call patterns account for the majority of spend.
- Review your system prompts for redundancy. Be ruthless. Every token in a system prompt is a token you pay for on every call that uses it.
- Ask whether each LLM call actually needs an LLM. Some tasks in your pipeline are probably deterministic logic dressed up as AI calls.
- Test smaller models on your lower-complexity tasks before assuming you need the largest model everywhere.
AI cost control isn't a one-time exercise. As your product evolves and your call patterns change, your cost profile will shift. Build the habit of reviewing it regularly, the same way you'd review any other operational cost that scales with usage.
The fundamentals of running a product responsibly haven't changed because the underlying technology is now a language model. You still need to understand where your money goes, make deliberate trade-offs, and design systems that can scale without the economics breaking. AI just raises the stakes for getting that right.