AI COGS Is a CEO Problem, Not a CFO One
For twenty years, gross margin was the most boring number on a SaaS P&L.
You built the software once, you served it from a rented server, and the marginal cost of the ten thousandth customer was close to the marginal cost of the first. Eighty percent was the benchmark. Anything below seventy-five raised an eyebrow. The number barely moved from quarter to quarter, which is exactly why nobody in the executive team spent much time on it.
That stability is over. Every AI request your product makes has a real, variable, usage-linked cost attached to it, and that cost lands directly in cost of revenue. The result is showing up in the benchmark data. Bessemer's State of AI 2025 puts LLM-native company gross margins at around 65 percent, well below the 80 to 90 percent that defined the previous decade of cloud software. ICONIQ Growth's State of AI research, which surveys roughly 300 software executives, has tracked average AI product gross margin moving from 41 percent in 2024 to around 52 percent in its 2026 snapshot. Improving, but nowhere near the old ceiling.
Here is the part that gets misfiled. Because gross margin appears on a financial statement, most companies treat AI COGS as a finance topic and route it to the CFO. Or, because the costs come from API calls, they treat it as an infrastructure topic and route it to the CTO.
Both routings are wrong, and the wrongness is expensive.
What AI COGS Actually Contains
When people say "our AI costs," they almost always mean the inference bill. That is the visible line, because it arrives as an invoice from a provider with a dashboard attached.
The full cost of serving an AI feature has at least five components:
- Inference. Tokens in, tokens out, across whichever models your product calls. The largest new line, and the only one most companies measure.
- Dedicated compute. Reserved GPU capacity, fine-tuning runs, self-hosted model serving. Fixed cost that behaves very differently from per-token spend and has to be amortised across usage that may not exist yet.
- The retrieval layer. Vector database reads and writes, embedding generation, document processing, storage and egress. Frequently a bigger surprise than the model bill itself, because it scales with corpus size and query volume simultaneously.
- Evaluation and observability. Quality monitoring, regression suites, LLM-as-judge scoring, tracing. Real spend that exists only because the product is probabilistic.
- AI-attributable support load. Tickets about outputs that were wrong, slow, or confusing. These resolve more slowly than conventional support tickets because reproducing the failure is harder.
Only the first of those five shows up on a provider dashboard. The other four are spread across cloud bills, SaaS subscriptions, and a headcount line, which is precisely why the total is so rarely assembled.
The Ownership Gap
Now ask a simple question: who in your company owns AI cost per unit of value delivered?
In most organisations the honest answer is nobody, and it is not because anyone is negligent. It is because the responsibility has been cleanly divided into three parts, each of which is being handled competently:
The CTO owns latency, reliability and output quality. Those are the metrics engineering is measured on, and the fastest way to improve all three is usually to reach for a larger model and a longer context window. Every one of those decisions is defensible on its own terms, and every one of them raises cost per request.
The CFO owns the invoice, the forecast and the classification. But finance receives spend as a monthly aggregate, weeks after the behaviour that caused it, with no dimension attached. You cannot manage a variable cost you only see in arrears and only in total. Finance can report the number. It cannot move it.
The product team owns adoption. They are rewarded for getting more users into the AI feature more often, which is the correct goal and also the mechanism by which COGS grows.
Cost per unit of value falls between all three. It is not on anyone's scorecard, so it moves without anyone deciding it should.
Only one person sits above all three functions. That is the structural reason this is a CEO problem, not something to delegate.
The Levers Are Not Finance Levers
There is a second, more practical reason.
There are really only four ways to fix an AI margin problem, and finance owns none of them outright:
- Price and package differently. Move from flat seat pricing to usage-linked or hybrid pricing, introduce fair-use ceilings, or reprice the tier where the heavy users sit. This is a go-to-market decision with revenue and churn consequences.
- Route models by task. Send the easy work to cheap models and reserve the expensive model for work that genuinely needs it. This is an engineering decision with quality consequences.
- Change the product surface. Cache aggressively, shorten context, batch, or redesign the interaction so it takes three calls instead of eleven. This is a product decision with UX consequences.
- Change who you sell to. Some segments are structurally unprofitable at your current price. This is a strategy decision.
Every one of those trades margin against something a different executive is measured on. Pricing versus growth. Model cost versus quality. Context length versus usefulness. Those are not calculations. They are trade-offs, and trade-offs between functions get resolved at the level above the functions, or they do not get resolved at all.
A CFO who spots the margin problem can raise it. Only the CEO can adjudicate it.
How the Erosion Hides
The most dangerous property of AI COGS is not its size. It is its shape. It moves slowly in aggregate while breaking badly in the detail, and there are three lags that keep it out of view.
The blend lag. Your reported gross margin is an average across all customers, including the ones who never touch the AI feature. Adoption at 30 percent hides economics that adoption at 90 percent will expose. The number you report today is a forecast of nothing.
The invoice lag. Behaviour changes on day 3. The invoice arrives on day 35. The board sees it on day 50. By the time the signal reaches the people who can act on it, the product has shipped two more releases.
The cohort lag. Usage in AI products is not normally distributed, it is heavy-tailed. A small group of power users can consume many multiples of the median. Those users tend to be your most engaged accounts, your best references, and your fastest expanders. The cohort you are least likely to look at critically is the cohort quietly setting your future margin.
Put those together and you get a company where the blended number drifts down by a point or two a quarter, which reads as noise, while the marginal economics of the customers you are actively trying to acquire have already inverted.
A Worked Example
Take a seat-based B2B product at $80 per seat per month. Conventional COGS of hosting, third-party services and support runs $16 per seat, giving the classic 80 percent gross margin.
You ship an AI assistant. Sixty percent of seats activate it. Among activated seats, the fully loaded cost of serving that assistant, inference plus retrieval plus evaluation plus the extra support it generates, comes to about $5.50 per seat per month.
Blended across all seats that is $3.30. Total COGS becomes $19.30 and gross margin lands at about 76 percent. Four points. Noticeable, not alarming, and easy to attribute to something else in a board pack.
Now run it forward eighteen months. Activation reaches 90 percent because the feature works and you promoted it. Usage per activated seat rises by half, because that is what happens when a feature becomes habitual. Fully loaded cost per activated seat is now $8.25, blended $7.40. Gross margin is roughly 71 percent.
Nine points below where you started, and still nothing that triggers an alarm in a monthly review.
Then look at the top decile of activated users, running about five times the average. That is roughly $40 of AI cost per seat per month against the same $80 of revenue. Add the $16 of conventional COGS and those seats carry a gross margin of about 30 percent. They are also the accounts your customer success team is holding up as proof the product works, and the ones your sales team is using as the template for the next hundred deals.
The blended number said you had a slow drift. The cohort number says your growth engine is aimed at your worst unit economics. Only one of those two numbers appears on a standard P&L.
Why You Cannot Fix This Retroactively
Here is the argument for acting before the number gets bad rather than after.
AI cost attribution cannot be backfilled. A provider invoice records that an API key spent a certain amount in a certain month. It does not record which customer, which feature, which environment, or which internal team generated that spend, and no amount of later analysis can reconstruct dimensions that were never captured at the moment of the request.
If you have shipped two quarters of AI features without per-request attribution, those two quarters are permanently unattributable. You cannot go back and ask what the enterprise tier cost to serve in Q2. The data to answer that question does not exist anywhere.
That is the difference between this and most reporting gaps. Usually you can rebuild the history from the source records. Here the source records are gone, and the only fix is to start capturing the dimensions now so that the next two quarters are answerable.
Which is a CEO decision, because it is a decision to spend engineering time on instrumentation before there is a visible crisis to justify it.
Questions Worth Asking This Week
None of these require a new system to ask. All of them require one to answer properly.
- What is our gross margin on the AI-enabled portion of the product, separated from the rest?
- What does it cost us to serve our most-used AI feature for one customer for one month?
- What is our cost to serve for the top decile of usage, and how does it compare to the median?
- How much of our AI spend is production traffic serving customers, versus development, testing and internal use? (Those belong in different places on the P&L, and mixing them makes production look worse and R&D look cheaper than either really is.)
- If usage doubled next quarter, what happens to gross margin, and is that number a calculation or a guess?
- Who is accountable for cost per unit of value, by name?
If the answers take a week to assemble, you do not have a cost problem yet. You have a visibility problem, which is the cheaper one to fix and the one that has to be fixed first.
This Is Not an Argument for Slowing Down
Nothing here says spend less on AI. Under-investing in AI capability is a considerably faster way to lose a market than serving it at 60 percent margin.
The argument is narrower: the cost of your AI product is now a strategic variable rather than an accounting outcome, and strategic variables need an owner, a target and a measurement cadence. Companies that can see cost per feature, per customer and per environment get to make deliberate trade-offs. They can price confidently, route models on evidence, and walk into a fundraise with defensible unit economics rather than an averaged number and an apology.
Companies that cannot see it will make the same trade-offs anyway. They will just make them by accident, and find out two quarters later.
Gross margin used to be the number nobody in the executive team argued about. In an AI-native business it is the number that determines your pricing model, your target segment, your infrastructure roadmap and your valuation multiple.
That makes it the CEO's number.
Further Reading
- AI FinOps: The Emerging Discipline Every CFO Needs to Understand: the operating model that sits underneath the questions above.
- How to Measure AI Cost per Feature (Not Just per Token): the practical mechanics of getting to a per-feature cost.
- Why AI Cost Attribution Is Broken (And How to Fix It): why provider dashboards cannot answer the cohort question.
- Is Your Internal AI Usage Distorting Product Profitability?: separating production COGS from R&D spend.
- The Real Cost of Model Sprawl (And How to Control It): what happens to COGS when model choice is ungoverned.
- Why Your AI Bill Explodes After Production: the usage curve behind the worked example.
Sources
- Bessemer Venture Partners, The State of AI 2025: LLM-native gross margins around 65 percent, with model costs running roughly 10 percent of revenue and 25 percent of total COGS.
- ICONIQ Growth, The State of AI: survey of roughly 300 software executives; average AI product gross margin tracked from 41 percent in 2024 to around 52 percent in the 2026 snapshot, with inference consuming a rising share of revenue as products mature.
- Development Corporate, AI COGS: The Hidden Margin Tax Every SaaS CEO Must Map Now: a complementary take on mapping the five AI COGS categories, and the piece that prompted this one.
- The SaaS CFO, Your AI Feature Is Quietly Destroying Your Gross Margin: on the per-seat arithmetic of adding an AI assistant to an existing plan.