Industry Insights 20 min read

Who Controls AI Application Cost Structure? Insights from the AI Supercycle Series

The article analyzes how inference costs dominate AI application economics, explains why pricing power lies with upstream model providers, examines Baseten's rapid growth, outlines three conditions that could shift cost control back to application companies, and highlights the sticky, asset‑heavy nature of inference infrastructure.

Fighter's World
Fighter's World
Fighter's World
Who Controls AI Application Cost Structure? Insights from the AI Supercycle Series

Key Insights

Inference costs overturn SaaS valuation logic. Inference is the largest variable cost for AI companies and scales linearly with users, ending the "build once, serve millions" era and invalidating 10‑20× SaaS revenue multiples.

Pricing power rests with an upstream competitor. About 90‑95% of inference spend flows to frontier models (OpenAI, Anthropic, Google), leaving AI application firms with little control over margins.

Cost‑control shift is already happening and is one‑way. Baseten processes 30 trillion tokens daily—more than OpenAI’s API—while migration away from frontier APIs is costly and rarely reversed.

Inference is extremely sticky. Downtime equals product downtime; migration requires re‑debugging the entire product, creating a strong moat.

Inference layer offers the highest profit certainty in the AI value chain. It is a utility‑like layer with heavy assets, economies of scale, and sticky pricing.

1. Inference Is the Main Business Cost, Not a Technical Overhead

Traditional SaaS incurs 5‑15% cloud hosting costs, with the rest being gross profit, supporting 70‑80% margins and 10‑20× revenue multiples. AI products differ: every user interaction consumes GPU time, HBM bandwidth, and power. Modern AI workflows involve hundreds of inference calls per user action, so more users and smarter products both raise costs rapidly.

"Inference is the COGS of AI value being delivered."

COGS (Cost of Goods Sold) ties directly to revenue; each dollar earned requires a proportional amount of inference compute.

2. Pricing Power Over Inference Costs Lies Upstream

Approximately 90‑95% of inference spend goes to frontier APIs, with only ~5% to custom models. Providers such as OpenAI and Anthropic set per‑million‑token prices, capping application‑layer margins.

Unlike SaaS where cloud providers compete, the frontier AI model market is an oligopoly (OpenAI, Anthropic, Google). Switching to a cheaper provider is often infeasible because model capabilities differ dramatically and migration costs are high.

Frontier providers also lock customers in by offering near‑free inference infrastructure in exchange for data, creating a “deep lock‑in” where migration costs become exponential.

"Before you know it, they’re post‑training models against those workflows that are sacred to you, that only you know."

3. Three Conditions That Could Transfer Cost‑Control Back to Applications

By 2026, three developments may enable application companies to regain control:

Condition 1: Open‑source models catch up. Open‑source models are ~90 days behind frontier models but can be run 70‑90% cheaper. Example: Cursor’s Composer 2.5 (based on Moonshot’s Kimi K2.5) scores 79.8% on SWE‑Bench Multilingual, comparable to Opus 4.7, at $0.50 per million tokens—1/5 to 1/10 of Frontier pricing.

Condition 2: Post‑training adds value. Baseten enables customers to bring data and utility functions, select a base model, and receive a highly tuned post‑trained model that often outperforms frontier models while costing ~70% less. Harvey’s legal agent benchmark shows a post‑trained model surpassing Opus 4.8 and GPT‑5.5 with a 0.913 pass rate.

Condition 3: Independent inference platforms provide full alternatives. Baseten’s all‑in inference stack offers optimized performance, multi‑cloud fault tolerance, observability, and integrated post‑training—far beyond raw GPU‑hour rentals from cloud providers.

4. Volume Shift: 30 Trillion Tokens per Day

Baseten’s daily token volume exceeds OpenAI API and Gemini combined, reaching ~30 trillion tokens. Revenue jumped from $200 M (Dec 2025) to $600 M (Mar 2026), a 1900% YoY increase, while peers like Together AI and Modal also saw rapid growth.

Customers now include not only AI‑native firms (Cursor, Poolside) but also traditional SaaS players (Notion, HubSpot, Clay, Superhuman), indicating a structural migration across the software industry.

5. Cost‑Control May Shift to Another Upstream

Even if applications regain some control, the next upstream—GPU supply—remains a bottleneck. Baseten’s GPU rental price doubled from $2.63/h to $5.10/h, reflecting soaring demand and limited supply.

GPU scarcity is structural; demand is continuous and growing, while chip capacity, data‑center construction, and power expand linearly.

Baseten estimates needing 150 k B200‑equivalent GPUs over two years (~$7 B compute spend). Building its own hardware reduces cost by ~30% versus renting, making heavy‑asset investment necessary.

Thus, while inference platforms may wrest cost‑control from frontier model providers, they inherit dependence on the GPU supply chain.

Conclusion: Inference Layer Holds the Highest Profit Certainty

Inference’s extreme stickiness creates three protective layers: high switching costs, price insensitivity, and a scale flywheel that lowers unit costs as adoption grows. This makes the inference layer more akin to a utility or refinery than a typical SaaS business, delivering durable, high‑certainty profits regardless of which models succeed.

Ultimately, the decisive factor for AI application economics is not which model is best, but who controls the final mile from silicon to intelligent services.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Cloud Computingindustry analysisAI economicsinference costGPU pricingBaseten
Fighter's World
Written by

Fighter's World

Live in the future, then build what's missing

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.