Token Costs Double Monthly, Yet 70% of AI Projects Just Please the Boss
Despite an 80% token price drop and a 20‑fold surge in global token usage, many enterprises face exploding monthly token bills, with 70% of AI projects merely placating managers; the article analyzes token economics, efficiency breakthroughs, and a task‑centric workflow that can halve costs and boost productivity.
01 Token Becomes the Biggest "Electric Bill" for Enterprises
In 2026, token prices for major models (xAI Grok 4.5, OpenAI GPT‑5.6, Meta Muse Spark) fell 80% within 24 hours, while global token demand surged 20‑fold year‑over‑year. China’s daily token calls reached 1.4 trillion, a >1,000× increase since early 2024, and China now accounts for 54.1% of global token consumption.
Bain & Company estimates that a company with 20,000 developers consuming $200 of AI tokens per developer each month incurs a $4.8 million token bill, with large enterprises seeing token costs double roughly every two months.
Cursor CTO Mike Saeks likened using the most powerful models for routine tasks to driving a Lamborghini to buy milk, highlighting the inefficiency of over‑powered models for simple work.
02 From "Cost per Million Tokens" to "Cost per Task"
Traditional pricing measured only per‑million‑token cost, but this metric is losing relevance. Kimi K3, priced at $3 per million input tokens and $15 per million output tokens, actually costs $0.94 per standard task—cheaper than lower‑priced competitors—because it uses fewer tokens and achieves higher success rates.
This shift prompted Artificial Analysis to formalize a "cost per task" metric in its Intelligence Index v4.1, moving the industry focus from token volume to task‑level value.
Palantir CEO Alex Karp called token‑based pricing "effing insane," arguing that enterprises pay for "worthless tokens" without a direct link to business value. Anthropic CEO Dario Amodei echoed this, noting that the same token can be worth a few cents in IT support but millions in drug‑design advice.
03 Full‑Stack Token Efficiency Revolution
Model layer: ByteDance’s ConceptMoE merges related tokens into higher‑level concepts, achieving 175% prefill and 117% decoding speedups, challenging the token‑as‑atomic‑unit paradigm.
Meta’s Muse Spark 1.1 uses "thought compression" to produce roughly half the tokens of competing models while maintaining comparable intelligence.
System layer: Intent Lab, founded by Jia Yangqing, boosted GLM‑5.2 inference speed from 102 tokens/s to 647 tokens/s—a 534% increase—reducing per‑token cost to about one‑sixth.
Infrastructure layer: The "Token Factory" concept emerged at WAIC 2026, with companies like SenseTime and PPIO delivering daily token services in the trillions and achieving 60‑90% KV cache hit rates. DeepSeek introduced peak‑off‑peak pricing, charging 2.4× more during high‑demand periods.
04 The Real Problem Lies in People, Not Tokens
Data from the Agentic AICon conference showed that 70% of enterprise AI projects exist merely to appease leadership, and 72.7% of decision‑makers view AI as a panacea. Large firms are restricting AI use because it becomes more expensive and slower.
Zhou Yu’s "SB AI Index" revealed that most project failures stem from flawed collaboration models rather than insufficient tokens or model strength.
Case studies illustrate this:
iOS app store submission: AI generated code passed tests but was rejected for violating Apple’s review rules, requiring two days of manual fixes.
Animation frame‑by‑frame pixel extraction: Feeding precise RGB values from original frames to the AI fixed the rendering in one attempt, saving dozens of AI iterations.
Control board design: After spending thousands of dollars on token‑intensive AI attempts, a $200 off‑the‑shelf board solved the problem instantly.
These examples underscore that human judgment remains the critical variable.
05 A Human‑Centric AI Development Process
Zhou Yu’s team (5‑6 engineers) produced 700 k lines of code, handling 30‑40 issues daily, by adopting a "task abstraction, not agent abstraction" methodology.
The workflow consists of:
Define tasks: Break vague requirements into clear, AI‑executable steps.
Value judgment: Assess whether AI outputs meet business goals.
All execution is delegated to AI; humans make only decisions. This is supported by a monorepo architecture and fully automated CI pipeline—tests, reviews, and deployments run automatically, giving the AI full context without needing to understand the entire project.
The result: 700 k lines of code from a small team, a realistic production system running for months.
06 First‑Principles of Token Economics
Supply‑side changes—price wars, new billing models, efficiency technologies, and token factories—make tokens cheaper and more abundant. However, without changing demand‑side usage, waste persists.
The key to saving tokens is ensuring each token delivers maximum value, as demonstrated by Zhou Yu’s team.
The industry is shifting from "how many tokens" to "how much value tokens create," echoed by Artificial Analysis, Palantir, and Anthropic.
07 Three Actionable Recommendations for Engineers
Adopt a "cost per task" mindset: Benchmark end‑to‑end cost for real tasks rather than per‑million‑token price.
Redesign collaboration, not just add agents: Assign tasks to AI and keep humans in the decision loop.
Secure inference infrastructure: Evaluate open‑source and on‑premise models to avoid vendor lock‑in and manage rising token spend.
Token costs are rising rapidly; treating them like electricity—cheap but abundant—requires disciplined, human‑centric processes to avoid waste.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Smart Era Software Development
Committed to openness and connectivity, we build frontline engineering capabilities in software, requirements, and platform engineering. By integrating digitalization, cloud computing, blockchain, new media and other hot tech topics, we create an efficient, cutting‑edge tech exchange platform and a diversified engineering ecosystem. Provides frontline news, summit updates, and practical sharing.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
