AI Is Compressing SaaS Gross Margins. Unless You Can See What Is Driving It.

Model selection, runaway usage, latency, quality, governance, and cost. Six AI problems, different owners, different tools. Five of them live inside the AI layer. One of them only looks like it does.

AI Is Compressing SaaS Gross Margins. Unless You Can See What Is Driving It.

For a long time, 80 percent gross margin was the gold standard of SaaS. Build the software once, sell it as many times as you like, and watch the margin hold because serving the next customer costs almost nothing. It was the business model that defined a generation of cloud companies and made software the most attractive asset class in venture capital.

That model is under pressure. For every $1 million in AI product revenue a SaaS company books in 2026, roughly $230,000 walks out the door as inference cost before a single engineer, salesperson, or marketer gets paid. That number comes from ICONIQ Growth's 2026 State of AI report, and it has quietly rewritten the rules of SaaS unit economics.

But the companies managing it well share one thing in common - they can see exactly what’s driving the cost. The ones who cannot are averaging it out and hoping the math works.

What Is Inference Cost?

Before going further, let's define this. After all, it’s at the center of this conversation.

Inference cost is the price of the AI doing its work. Unlike traditional software that runs the same code for free once it’s built, every AI response requires computation, billed at a fraction of a penny. That fraction adds up fast when thousands of users are triggering AI features dozens of times a day.

What the Numbers Show

The data from 2025 through early 2026 is consistent across multiple sources.

Traditional SaaS companies historically targeted gross margins of 80 to 90 percent. The logic was simple: once the software was built, the marginal cost of serving one more customer was trivial. A few more server cycles, a bit more storage. Rounding errors.

AI changes that logic at the root. Every time a user triggers an AI feature, the SaaS company makes a real-time call to an AI model provider and pays for the AI's work by the token. The cost does not amortize across customers. It does not decline as the user base grows. Every interaction, every prompt, every response is an additional cost.

Bessemer Venture Partners' February 2026 pricing playbook puts AI-native company gross margins at 50 to 60 percent. ICONIQ's 2026 State of AI puts the average AI product gross margin at roughly 52 percent. Across Q4 2025 and Q1 2026 earnings seasons, a new operating corridor of 60 to 70 percent has established itself among publicly listed SaaS providers who openly discuss AI-driven margin pressure.

Several public SaaS companies disclosed 6 to 9 percentage points of year-over-year gross margin compression in Q4 2025, with explicit attribution to AI feature cost. These are not small companies making experimental bets. These are established SaaS businesses watching a core financial metric move in the wrong direction.

How It Happened So Quietly

A large part of why AI spend caught so many companies off guard is that the early stages were cheap and predictable.

Many SaaS companies started with flat-rate API subscriptions or bundled enterprise agreements. The bill was fixed. Teams built features, users adopted them, and the cost was invisible. Then usage scaled, those agreements ended, and companies moved to consumption-based pricing. Suddenly the cost model changed completely. The visibility to catch it early did not exist.

This is not a story about recklessness. It is a story about a cost structure that behaves differently from anything SaaS companies have managed before. Traditional infrastructure costs scale gradually and relatively predictably. AI inference costs are dictated by user behavior, which is variable, hard to forecast, and invisible at the customer level until the bill arrives. Currently invisible, that is.

The CFO who asks "why did our gross margin drop four points this quarter" deserves a real answer. In most companies right now, the honest answer is: we do not know which customers, which features, or which workflows drove it.

What Investors Are Starting to Ask

The financial disclosure patterns from Q1 2026 are worth paying attention to. Several public SaaS companies began disclosing AI inference cost ratios separately in their financial filings, typically 4 to 9 percent of revenue. The companies that disclosed received analyst credit for transparency. The companies that bundled inference costs into broader infrastructure categories drew skeptical questions.

That dynamic will filter down to private companies. PE sponsors and VC investors who see public companies being rewarded for AI cost transparency will ask the same questions of their portfolio companies. The CFO who can answer "our AI inference cost is X percent of revenue, and here is our cost per customer by segment" will be in a very different conversation than the one who cannot.

This is not a distant scenario. It’s happening now at the public company level and moving down the market.

The Companies That Will Hold Their Margin

Not every SaaS company with AI features is watching its margin collapse. Remember a few paragraphs earlier I said currently invisible at the customer level? It doesn’t have to be. The SaaS companies holding their ground share a common characteristic: they can see their costs at the customer level and act on what they see.

Model routing is part of it. The well-run companies have learned to route simple queries to cheaper, faster models and reserve the expensive frontier models for complex tasks. That alone can cut inference costs by 50 to 70 percent without degrading the user experience.

Caching is part of it. Both Anthropic and OpenAI now offer significant discounts on cached input tokens. Products with stable system prompts and repeated context windows can cut effective per-query costs substantially with relatively modest engineering effort.

But neither of those optimizations is possible without first knowing which customers, which features, and which workflows are driving the cost. Optimization without attribution is guesswork. And guesswork does not hold a margin.

The Strategic Question

The SaaS companies that built the 80 percent gross margin model did it by understanding their costs precisely and pricing accordingly. The economics were visible. The margin was defensible.

AI has changed that. It has introduced a cost layer that’s variable, customer-driven, and invisible at the level of granularity required to manage it. The bill arrives. The explanation does not.

The most important question SaaS companies need to ask right now is not "how do we cut AI costs" but "do we actually know which customers, features, and workflows are driving them?" Because the margin compression you are experiencing may not be structural and unavoidable. It may be an attribution problem masquerading as an economics problem.

Those are very different situations with very different solutions. And you cannot tell which one you are in until you can see what is actually driving your costs.

Request a demo

About the Author

Photo of Alan Cox
25+
Years Experience
Alan Cox

CEO and Co-Founder

Leadership Team

Alan Cox founded Beakpoint after experiencing firsthand the frustration that comes with mysterious cloud costs. As a technology leader who has spent over two decades building and scaling software organizations, he's seen how cloud expenses can spiral out of control.

Expertise

strategy
leadership
cost accounting
software engineering
cloud operations
aws
+2 more

Previously at

Geoforce (VP of Software Engineering)SignalPath (CTO)

Related Articles