AI ECONOMICS
How Tokens Became the Unit of Production of the AI Economy
AI costs are doing something strange right now.
The price of each token keeps falling, but company bills keep climbing anyway.
Enterprise AI budgets jumped from $1.2 million in 2024 to $7 million in 2026. 73% of companies blew past their AI cost projections last year.
Inference, the cost of actually running AI models, now eats up 80 to 85% of total AI spend.
AI agents do not send one request and stop. They plan, check, retry, and call other tools, sometimes 10 to 20 times for a single task. Background AI agents also run around the clock, watching documents and systems even when no one asked them to.
Each of these steps burns tokens, and the bills add up fast.
There is good news too.
Companies that route simple tasks to cheaper, smaller models and save the expensive ones for hard problems are seeing costs 8 times lower than companies that send everything to the priciest model.
FinOps teams are stepping up fast here.
- Only 31% managed AI spend in 2025.
- Now 98% do.
One more warning worth noting. Some experts believe today's AI prices are being kept artificially low and will rise later.
Building cost models and system designs based on today's cheap prices is risky.
The main lesson for any finance or tech leader: build a clear budget and tracking system for AI spend now, route work to the cheapest model that gets the job done, and plan for prices to change before they actually do.
BEST PRACTICES
Governing Gemini API Costs Before They Happen: Google API Quotas as a FinOps Guardrail
Generative AI costs can spike fast. A single mistake, like a stuck script or a leaked API key, can turn a normal bill into a five-figure shock in just a few hours.
This is why Google Cloud API quotas deserve more attention from FinOps teams.
- Quotas cap how many requests a model can handle each day.
- They stop runaway usage before it turns into a runaway bill.
- Billing alerts only tell you after money is already spent.
- Quotas act earlier, more like a speed limit than a fuel gauge.
Google also just introduced Spend Caps, now in private preview. It pauses API traffic once you hit your budget, but keeps your systems running. You can turn it back on anytime.
Quotas and Spend Caps are not competing tools. Quotas control which models get used and how much. Spend Caps control the total dollar limit.
Together they give teams real guardrails, not just visibility.
A new open-source tool called gemini-quota-manager makes this easier. It lets teams review and update quotas safely, with dry-run checks built in by default.
Cost control for AI needs more than watching a dashboard. Real limits, set before spending happens, are what keep budgets safe.
EVENTS
Webinars
Discover how to measure the real business value of AI, optimize cloud spend, and prove the ROI of your AI initiatives with practical FinOps strategies and expert insights.
📅 July 16
🕚 6:00 PM CEST (Spain) / 12:00 AM EST
Meetups

Our next meetup will explore a topic many organizations are already facing: how to integrate AI, automation, and new operating models without losing control, efficiency, or visibility.
This session is designed for professionals working in Cloud, FinOps, AI, Digital Transformation, and Operations. Expect technical talks, engaging discussions, and networking with the community.
October 7 – Madrid (Utopicus Nuevos Ministerios)
https://luma.com/036wqk98
AI COST ANALYSIS
Understanding Token Cost Anatomy
AI spending is growing fast, and most invoices only show a total.
They do not show why the bill got so big. Input tokens make up the largest share, around 38 percent.
These come from system prompts, chat history, and retrieved documents.
The biggest fix here is prompt caching, which can cut costs by up to 90 percent if your prompts stay the same across calls. Output tokens cost even more per token than input.
Setting a hard limit on response length can save real money at scale, sometimes tens of thousands of dollars a year from a small engineering fix.
Embedding tokens are cheap alone, but wasteful when documents get re-processed for no reason.
Fine-tuning costs look big and obvious, but the real cost often hides in ongoing inference bills after training.
Inference infrastructure costs pile up when GPU capacity sits idle and unused.
None of these problems can be fixed without visibility first.
Teams that tag and track spend from day one have a real shot at controlling AI costs. Teams that wait to look only after the bill arrives will keep guessing.
Seeing where AI dollars go is the first step to controlling them.
🎖️MENTION OF HONOUR
Best AI Gateways for Cost Tracking and Chargeback
AI spending is quickly becoming a real headache for finance and engineering teams. As more companies use AI models like GPT and Gemini, tracking who spent what gets messy fast.
These gateways sit between your apps and AI providers. They track usage, set spending limits, and help bill the right teams for what they use.
Here are the key players compared:
- Bifrost gives the most complete control.
- It lets teams set budgets, track spending by project or person, and create audit logs for compliance.
- It also supports exporting data to tools like BigQuery, which makes it useful for larger companies.
LiteLLM is a open source option.
It works well for small teams that just need basic spending caps. OpenRouter offers one bill for many AI providers. This makes accounting easier but offers less detailed cost control.
Cloudflare AI Gateway focuses on speed and caching. It helps cut costs indirectly by reducing repeat requests, though it lacks direct billing tools.
The takeaway is simple.
As AI use grows, so does the need for clear cost tracking.
Picking the right gateway can mean the difference between predictable AI budgets and nasty billing surprises.
Save 20% on AI Value & FinOps Certifications

The job market is hungry for certified professionals who can prove results. Don't let your company's budget leak due to a lack of specialization.
Use code: FINOPSWEEKLY_20 to get an instant 20% discount on the most prestigious certification bundles:
FinOps for AI
FinOps Certified Practitioner
FinOps Certified Engineer
FinOps Certified FOCUS Analyst







