TOGETHER WITH CLOUDZERO
Where Does Your AI Spend Governance Land?
This 2-minute benchmark asks five questions about how you govern AI spend, then shows exactly where companies like yours land against 260 senior finance leaders.
What you'll see:
Your cohort, and the peers who avoid the damage
The one variable that separates them
Your likely exposure in dollars, every stat with its sample size
TOKEN ECONOMICS
The Token Bill Is the Cheapest Number in Your AI Budget
Most companies are watching the wrong number when it comes to AI costs. The token bill looks scary on the invoice, but it is actually the cheapest and fastest-falling cost in the whole system.
Token prices keep dropping. Some Asian AI providers cut prices by 50 to 99 percent in a single quarter. But total spend keeps climbing anyway.
Why? Volume, not price, is the real driver.
Agentic AI workflows use 5 to 30 times more tokens per task than a simple chatbot question. One customer service task went from 4 cents to 1.20 dollars in three years, even as prices fell the whole time.
Research backs this up.
- MIT found 95 percent of AI pilots show no real business impact.
- IBM found only 25 percent of AI projects meet ROI goals.
- Gartner expects 40 percent of agentic AI projects to be canceled by 2027.
The real problem is ownership. Someone needs to own model choices, agent limits, and design decisions before deployment, not after the bill arrives.
New EU rules starting in August 2026 make this even more urgent, with fines up to 7 percent of global revenue. Dashboards help track spend. They do not fix the decisions that created it.
Watch the architecture, not just the invoice.
AI PROVIDERS
Cost & Pricing Updates of AI Providers

OpenAI
- Organization and project-level spend limits are now live for the API platform, letting teams set monthly caps to monitor spend or fully block API calls once a limit is hit. This gives teams a built-in guardrail against billing surprises without manual tracking.
Anthropic
- Claude Code v2.1.218 fixes Bedrock spend metering, so application-inference-profile ARNs and other config-mapped model IDs are now billed at the correct configured rates. This fixes silent cost-tracking errors for Bedrock inference profile users.
- Server-managed settings no longer trigger approval prompts for benign feature and cost toggles, making it easier to adjust cost-related settings in managed environments.
- Claude Code v2.1.216 fixes double-counted session costs caused by multiple cumulative message_delta frames, so session cost and token numbers now match actual spend.
- The same release fixes a long-session slowdown from quadratic message normalization, plus metrics endpoint formatting bugs, improving both performance and cost data reliability for long sessions.
AI COST OPTIMIZATION
Building LLM Observer: an open MVP for identity-aware LLM FinOps
AI tools are getting easier to use, but tracking who is spending what on them is still hard. A developer built an open source tool called LLM Observer to solve this exact problem.
It watches API traffic from AI models like GPT or Claude and shows which user, team, or app is generating costs.
It connects to company directories like Entra ID so spending gets tied to real people and departments, not just a provider invoice total.
It flags waste automatically, things like repeated large prompts with no caching, expensive models used for simple tasks, or too many retries.
Each problem comes with a suggested fix and an estimated savings amount. Teams can also plug in their own negotiated pricing instead of relying on public rate cards.
This is not a finished commercial product.
The creator is upfront that it has not been tested on a live cloud deployment or connected to a real company directory yet.
It is also not meant to replace bigger platforms like Langfuse or LiteLLM.
Cost visibility for AI spending does not need to be complicated, it needs the right metadata and a little curiosity.
WEBINAR
FINOPS FOR AI: See It. Control it. Use it Safely.
Join this FinOps for AI webinar to learn how to control AI costs, improve visibility, implement governance, and justify AI investments with confidence.
📅 August 20, 2026
🕚 6:00 PM Spain / 12:00 PM ET
VIDEOS & PODCASTS
GCP AI Cost Optimization: Vertex AI Model Selection Guide
How to Choose the Right AI Model on GCP Without Blowing Your Budget. We sit down with Vera to break down how to select the best AI model on Google Cloud's Vertex AI, balancing pricing, performance, and business use cases.
AI GOVERNANCE
Enterprise AI spend and cost management
AI spending is growing fast, and most companies cannot keep up. A McKinsey survey found that 93 percent of organizations are going over their AI budgets.
Many teams do not know how much they are spending on AI because purchases happen across different departments and tools. Employees are also building AI agents and workflows on their own, which adds hidden costs that are hard to track.
AI costs behave differently than normal tech spending.
The same task can use very different amounts of computing power depending on the tools and steps involved. This makes it hard to predict costs or set clear budgets.
McKinsey suggests treating this like FinOps for AI. Companies that manage this well can save 20 to 30 percent on AI costs.
The key steps include tracking spending in one place, matching costs to real business value, choosing the right model for each task, and building flexible sourcing strategies instead of locking into one vendor.
Automation and built-in guardrails can help make cost efficient choices the default, not an afterthought.
AI can create huge value, but only if leaders build strong habits to track, manage, and question how it is used.
Companies that build this discipline early will control costs and use AI more wisely than those still finding out the price after the damage is done.
AI COST OPTIMIZATION
AI cost controls are coming. UX needs to make sure users do not pay the hidden price.
A recent McKinsey report found that 93% of companies went over their AI budgets. Many businesses are now adding cost controls, like cheaper models, shorter answers, and usage limits.
These choices sound like a finance or tech fix, but they change how people actually experience AI at work.
If a chatbot forgets details or gives a shallow answer, the person using it feels the cut, even if they never see the cost decision behind it.
This can push employees back to unofficial tools, since the approved system feels slower or less helpful.
A cheaper AI answer is not always a cheaper outcome.
If a customer has to repeat themselves or call a human agent, the saved cost just shows up somewhere else, like more staff time or lower satisfaction.
Cost control needs design input from the start, not just as a warning message added later. The real measure of success is not cost per token.
It is the cost per successful, trustworthy result, counting both AI spend and the human time needed to fix or complete the job.
Cutting AI costs without watching the full picture can quietly raise the cost of trust and service quality.
🎖️MENTION OF HONOUR
14 Billion Tokens Later: AI Cost Optimization Is a Systems Design Problem
This article makes a simple but important point for FinOps teams: AI costs are not really about token prices. They are about how your workflows are built.
The writer once used 14.6 billion tokens. That number sounds big, but it does not tell you how many tasks actually got done well. He points out four hidden cost drivers that FinOps teams should watch closely.
Systems often resend the same background information over and over, quietly padding costs.
Teams use expensive, powerful models for simple tasks that a cheaper tool could handle just fine.
AI agents can loop and retry without limits, quietly burning budget while looking "busy."
When AI output needs heavy human editing, the real cost has not gone down, it has just moved to a person's time.
His suggested fix is to track cost per accepted outcome, not cost per token.
He also favors hybrid setups, where local tools handle prep work and cloud models only step in for harder tasks.
The main message is this: good AI cost control starts with smart workflow design, long before the invoice arrives.
Save 20% on AI Value & FinOps Certifications

The job market is hungry for certified professionals who can prove results. Don't let your company's budget leak due to a lack of specialization.
Use code: FINOPSWEEKLY_20 to get an instant 20% discount on the most prestigious certification bundles:
FinOps for AI
FinOps Certified Practitioner
FinOps Certified Engineer
FinOps Certified FOCUS Analyst











