What it costs to run AI, and where the money leaks
By Sara · · 6 min read
A website costs the same whether ten people visit or ten thousand read every page. An AI feature does not. It bills by the word — every word that goes in, and every word that comes back — which is why the first invoice is a pleasant surprise and the fourth one is not. The useful part: most of that bill is not work. It is repetition.

Why the bill grows when nothing changed
An assistant has no memory between messages. To answer the fifth question in a conversation, it is handed the whole thing again from the top: who it is, how it should behave, what your business sells, and everything already said. Then the sixth question arrives and it all goes again, slightly longer.
So the cost of a conversation does not grow in a straight line. It grows with the square of how long the conversation runs. A chat that goes twenty messages deep costs far more than four chats of five messages, for the same amount of actual help. That single fact explains most surprise invoices.
Four leaks worth fixing
None of these are about writing shorter prompts or accepting worse answers. They are about not paying for the same thing repeatedly.
- Paying twice for the same words. Your instructions, your price list, your policies — the fixed part of every request — are sent again with every single message. Sent the same way each time, that repeated part can be served from a cache at about a tenth of the price. Published results from teams who fixed this report input costs falling by 70 to 88 per cent. The catch is order: anything that changes — today’s date, the customer’s question — has to come last, or the saving is lost.
- Carrying the whole conversation forever. An assistant re-reads the entire conversation to answer the next message, including the steps it finished half an hour ago. Clearing what is done keeps it cheap. In published tests of a long task, clearing what is finished cut the words used by about 84 per cent — and the answers got better, not worse, because there was less clutter to wade through.
- Reading the whole toolbox before every answer. Connect an assistant to your systems — orders, calendar, invoices, WhatsApp — and the description of every one of those tools is read before it does anything at all. A busy setup can spend tens of thousands of words on that menu before the first word of real work. Loading the two or three tools a request actually needs cuts most of it, and the assistant also picks the right tool more often.
- Rush pricing for work that can wait. Tonight’s summaries, tomorrow’s listings, this month’s classification — none of it needs an answer in two seconds. Work submitted as a batch, to be done within the day, costs half. Most businesses are paying the rush rate for everything because nobody separated the two.
The one check worth running this week
Ask whoever runs your AI one question: what share of what we send is being served from cache? There is a number for it, and it takes a minute to look up. If it comes back at zero, you are paying the full rate for the same words on every request — and that is usually a day’s work to fix, not a rebuild.
Three quieter savings
- Right-size the thinking.
- The same model can be told to think harder or less hard. Routine jobs rarely need the deep setting, and turning it down is usually a better trade than switching to a weaker model — which quietly throws away the savings in the first point above.
- Keep the answers to questions you are asked constantly.
- “What are your timings?” does not need to be worked out from scratch every time. Storing answers to common questions saves both halves of the bill, not just the question half — and for a support or enquiry line, repeat questions are often 20 to 45 per cent of everything asked.
- Send the heavy reading somewhere else.
- When a job means reading fifty documents to produce one paragraph, the reading belongs in a side task that hands back the paragraph. The main conversation stays short, and short is cheap.
The number to judge it by
Not the price per request. The price per finished job — per enquiry answered, per invoice read, per listing published. A cheaper setting that needs three attempts to get the job done is not cheaper, and a bigger model that gets it right first time is often the lower bill at the end of the month.
Set it against the thing the AI replaced. An enquiry line that answers at eleven at night for a few rupees a conversation is not competing with a free alternative — it is competing with a missed enquiry.
Where Digitalgrub is in this
We run this software ourselves, not just build it for other people, so the running cost is our problem too. Our Tamil Nadu jobs site rebuilds itself and publishes every day without anyone pressing publish, and it costs under a rupee a day to keep running. That number is not luck — it is the four points above, applied to a job that repeats forever.
When we build an AI feature for a client, the running cost is part of the design rather than a surprise at the end. We measure what the thing actually spends before changing anything, we fix the repetition before touching quality, and when a saving would cost you accuracy, we say so and let you decide instead of quietly making the answers worse.
Working with us
If you are already paying an AI bill and it is drifting upward, it is worth a look — the leaks above are common, and most of them are structural rather than clever. If you are planning an AI feature and want to know what it will cost to run before you commit, tell us what you have in mind and we will give you a straight answer.
Talk to us about your AI bill