This website uses cookies

Read our Privacy policy and Terms of use for more information.

Presented by

Want to appear here? Talk with us

TOGETHER WITH OUR SPONSOR
AI Just Got Its Last Blank Check.

AI introduces a new cost model. Every team can contribute to token spend.

This guide shows you how to apply FinOps principles to AI. Prioritize your efforts and build the visibility, governance, and accountability needed to scale AI with confidence.

AI ECONOMICS
Why Cost Per Token Is the Wrong AI Metric

Every AI model looks cheap or expensive based on one number: cost per token. But the real cost is cost per successful task..

A budget AI model might cost far less per token. But if it fails often, your engineers spend expensive hours fixing broken code or bad output. That hidden repair cost can make a "cheap" model far pricier than a premium one. A real example:

A frontier model costs 13 times more per token, but ends up nearly 5 times cheaper once you count the human fixes saved. The break-even point depends on your team's pay rate.

High-cost engineering teams gain value from premium models almost immediately. Lower-cost teams only benefit when tasks are genuinely hard.

The smart move is not choosing one model for everything. Route simple, low-risk tasks to cheap models. Send complex, high-stakes work to stronger models.

For FinOps teams, this means AI budgets should track task success, not just token bills. The invoice tells you what you paid. The rework tells you what it really cost.

AI PROVIDERS
Cost & Pricing Updates of AI Providers

OpenAI

  • GPT-5.6 launches with cost-optimized models: Terra (balanced intelligence) and Luna (high-volume efficiency). New explicit prompt caching controls and persisted reasoning reduce redundant computation, empowering FinOps teams with leaner, predictable token spend across high-volume API workloads.

Anthropic

  • A critical Claude Code regression was patched, fixing a bug that bypassed prompt caching and overcharged Bedrock, Vertex, Mantle, and Foundry users for trailing system context. This resolves unexplained invoice cost creep by ensuring tokens are properly discounted.

  • Additionally, Claude Code reduced CPU, memory, and transcript overhead while fixing its session cost counter and usage reporting. This restores accurate spend observability and reliable budget tracking for engineering teams.

AI COST ALLOCATION
I Set Up Per-Team LLM Cost Tracking for Amazon Bedrock in 30 Minutes

Shared AWS accounts make Amazon Bedrock costs a mystery. Multiple teams use the same account, but the bill just shows one big number. No one can tell which team is spending the most, or who is using pricey models when a cheaper one would work fine.

A free, open-source tool called AgentGateway sits between your apps and Bedrock. It tracks every request by team, counts tokens, and calculates the real dollar cost.

Teams get their own key, and a built-in dashboard shows spending by user, team, or model.

In a test, one team using a powerful model called Claude Opus drove most of the total spend, even though it made up only a small share of all requests.

That kind of insight is exactly what FinOps teams need for budget talks and chargeback decisions.

The article also points out common setup mistakes, like using the wrong model ID format or wrong pricing labels, that cause costs to show as zero.

For companies running AI on shared cloud accounts, this kind of visibility turns a confusing bill into clear, team-level answers.

WEBINAR
FINOPS FOR AI: See It. Control it. Use it Safely.

Join this FinOps for AI webinar to learn how to control AI costs, improve visibility, implement governance, and justify AI investments with confidence.

📅 August 20, 2026
🕚 6:00 PM Spain / 12:00 PM ET

EVENTS
The Return of the Best Online FinOps Event

Join FinOps professionals at the FinOps Weekly Summit 2026 and discover how to:

  • Transform FinOps from reactive cost control into strategic business value

  • Master AI-driven cloud and infrastructure optimization

  • Build scalable unit economics for the AI era

  • Learn proven strategies from organizations managing billions in cloud spend

Limited seats available.

📅 October 20 & 21, 2026

AI COST ANALYSIS
Why Your LLM Bill Is 3 What the Pricing Page Promised - DEV Community

If your LLM bill keeps coming in three times higher than the pricing page promised, you are not alone. There are five hidden cost leaks between the sticker price and your actual invoice.

Output tokens cost three to five times more than input tokens, and your specific use case decides how much this hurts.

Chat costs less, but translation and code generation can cost nearly three times more, even on the same model.

Each provider counts tokens differently, so a cheaper price per token can end up costing more once your real text is measured.

Prompt caching offers discounts up to 90 percent, but most teams never turn it on, leaving real money on the table.

Batch processing gives you half off for tasks that do not need instant answers, like nightly reports or data labeling.

Failed requests and retries quietly charge you twice, and the fix is simple backoff and better monitoring.

Stacked together, these leaks explain a 40 to 65 percent gap between the quoted price and the real cost.

Measure your real usage, turn on caching, and route traffic wisely, because small configuration choices can save tens of thousands of dollars a year.

LLM COST MANAGEMENT
I was using a frontier model to update jira tickets. here's how i fixed it.

Every FinOps team should read this story about a developer who found something surprising in his AI bill.

Don't Burn Your AI Budget: I'm Routing AI Models Like Infrastructure Last month I looked... DEV Community

He was using an expensive frontier AI model for every task. That included small jobs like fixing Jira tickets and big jobs like writing architecture plans.

So he built a simple system with three cost tiers. Small, cheap tasks go to free or low-cost models. Medium tasks go to a mid-priced model.

Only hard, high-value work like architecture and security reviews goes to the expensive model.

Routing work to the right tier is not the same as handling outages. Mixing these two ideas made his cost data confusing and unreliable. He found that the biggest cost driver was not the model itself.

It was how the system managed retries, context, and failures. A messy process made every model more expensive. A clean process made even costly models cheaper.

This piece is a strong reminder for FinOps leaders.

Cost control in AI systems depends on smart workflow design, not just model choice. Matching the task to the right tool can save real money without slowing down teams.

🎖️MENTION OF HONOUR
The Metric CFOs Struggle to Track: AI Usage - WSJ

AI costs are starting to look a lot like an unpredictable utility bill. As AI vendors move from flat fees to charging by the token, finance teams are struggling to keep up.

A new KPMG survey found only 26% of companies have a full picture of their AI spending.

Half have partial visibility, and the rest find out costs only after the bill arrives. Some companies have seen token usage jump sixfold in months, blowing past annual budgets.

Here is how different companies are handling it.

- Affirm tracks token usage in near real time and reviews it weekly with leadership. Corning limited which AI tools employees can use, while still letting staff experiment daily.

- Reckitt found that adoption of some AI tools dropped off after a few weeks and slowed a rollout when data quality suffered.

- Amer Sports is moving slowly on AI projects to avoid setting up tools without clear long-term value.

Analysts see echoes of the cloud spending boom during the pandemic, when companies overspent on software and later cut back.

The message for finance teams is simple: Track AI spending closely now, or risk a costly surprise later.

Save 20% on AI Value & FinOps Certifications

The job market is hungry for certified professionals who can prove results. Don't let your company's budget leak due to a lack of specialization.

Use code: FINOPSWEEKLY_20 to get an instant 20% discount on the most prestigious certification bundles:

  • FinOps for AI

  • FinOps Certified Practitioner

  • FinOps Certified Engineer

  • FinOps Certified FOCUS Analyst

FinOps Weekly

FinOps Weekly

Save on Your Cloud Costs with 5 Minutes every Sunday

Liked the Newsletter?

Share your thoughts!

Login or Subscribe to participate

Keep Reading