Writing

AI Models Are the New Fiber

By Jay Zeng | August 3, 2026


I’ve been building software for 15 years. I have never seen an industry work this hard to destroy its own pricing power.

Here is what it costs to process one million tokens across the major AI providers today:

Model Input (per 1M tokens) Output (per 1M tokens)
DeepSeek V4 Flash $0.14 $0.28
DeepSeek V4 Pro $0.44 $0.87
GLM-5.2 ~$0.60 ~$3.00
Kimi K3 $3.00 $15.00
GPT-5.6 ~$5.00 ~$25.00
Claude Fable 5 ~$10.00 ~$50.00

DeepSeek V4 Flash costs fourteen cents per million input tokens. That is not a typo. And the weights are open — MIT license — so you can self-host and pay zero per token if you want.

Claude Fable 5 costs roughly $50 per million output tokens. Same task, same output, 178 times the price.

What exactly is that premium buying?

The Switch Is Already On

In January 2025, Chinese AI models accounted for less than 5% of US enterprise token usage. By April 2026: 46%. DeepSeek has been #1 on OpenRouter for seven consecutive weeks at 17.6% market share. Chinese models collectively hold 44% of token volume across the top 10 (platform-wide: 53%). By July, the figure hit 58%.

These are production workloads, not experiments.

Lindy, a 25-person AI startup in San Francisco, migrated 100% of its traffic from Claude to DeepSeek V4. CEO Flo Crivello said it was survival — their API bill had exceeded total employee salaries. The result: inference costs dropped 90 to 95 percent. Performance held.

Coinbase set GLM-5.2 and Kimi as the default models for its engineering team. Brian Armstrong announced it publicly. AI spending was cut by half.

Enterprise spend data from Ramp shows DeepSeek topping the trending vendor list in June 2026 — not “trending AI,” trending among all software vendors. Companies are paying DeepSeek directly for hosted API access. This is not hobbyists chasing a discount at the margins. The core workloads are moving.

The price curve only bends one direction. It keeps going down.


What Moat?

Every time I hear the argument that US model companies are defensible, it comes down to one of three claims. Let me take them in turn.

Claim 1: “Our models are better.”

True — for now. But the gap is collapsing.

When your model is 1.04× better at 178× the price, the math breaks for anyone except the most latency-sensitive or regulated workloads. And that slice of the market shrinks every quarter.

Claude is the best coding model today. The question is whether “best by a whisker” sustains a 178× premium. Claude Code is free software and model-agnostic — you can plug in any provider by changing one line of config. Developers use it with Claude models because Claude produces the best code. That is the only reason. The moment benchmarks converge, the switch is trivial.

A moat made of one benchmark, shrinking every quarter, is not a moat.

Claim 2: “We have cloud distribution.”

This one sounds convincing until you check.

DeepSeek has been on AWS Bedrock since March 2025, when R1 became the first DeepSeek model offered as a fully managed, generally available service. 1 V3.1 followed in September 2025, and V3.2 was added in February 2026. All three are listed as “Amazon sold models,” available immediately, no marketplace subscription required. Same AWS console. Same one-click deployment. Same consolidated bill.

It gets worse. Qwen is available on Bedrock and Azure AI Foundry. GLM-5 is on Google Vertex AI. AWS added 18 open-weight models in April 2026 alone. They will keep adding whatever their customers demand.

Cloud providers do not care which model you run. They make money on compute — training, serving, embeddings, vector search. The model is a SKU. SKUs are interchangeable.

This is not speculation. Watch the incentives. Amazon’s Bedrock AI revenue has gone from 9% of AWS AI revenue to 37% in one year, and 80 to 90 percent of Bedrock customers use Claude. But Amazon’s 10-year, $100 billion-plus compute commitment with Anthropic is about locking in compute spend — not about model exclusivity. The moment enough Bedrock customers ask for DeepSeek V4, it will be there. Amazon sells compute. Model loyalty is someone else’s problem.

Now look at what Anthropic pays for this “exclusivity.” An estimated $1.9 billion in cloud partner payouts in 2026, projected to hit $6.4 billion by 2027. And there is an accounting note worth reading: Anthropic reports revenue on a gross basis — the cloud partner’s cut is included in the headline number. OpenAI reports net, deducting Microsoft’s share first. Part of that $10.9 billion quarterly revenue figure belongs to Amazon, Google, and Microsoft.

Cloud distribution was never a moat. It is a toll road. And toll roads charge by the mile.

Claim 3: “Geopolitics protects us.”

This one is real — but not for the reason most people assume.

The data-sovereignty argument is already obsolete. DeepSeek runs on AWS Bedrock. Qwen runs on Azure. These models sit on US infrastructure, processed by US servers. No data leaves American soil. If the concern was where your bytes go, it is already solved.

The actual geopolitical question is harder, and simpler: will US policymakers ban Chinese-origin models outright, regardless of where they run? Not because of a technical vulnerability — because of the name on the label.

If that happens, it is a genuine moat for US labs — the strongest one they have. But it is a moat built by policymakers, not by product teams. It can be expanded. It can be narrowed. It can be carved out with exceptions. Companies that rely on it are not competing on what they build. They are competing on lobbying.

And here is what makes this fragile: the 46% of US token traffic already flowing to Chinese models represents real businesses making real economic decisions. Every one of those companies has a lobbyist. Every one will argue that banning the tools that cut their costs by 90% is anti-competitive. Policy moats only hold until the economic pressure to breach them exceeds the political cost of maintaining them.

A moat you do not control does not belong to you. It belongs to whoever does. And they can take it away — or give it to someone else.

Cloud distribution was never a moat. It is a toll road. And toll roads charge by the mile.

Model Quality Cloud Distribution Geopolitics
Anthropic has it? Yes (narrowing) Already breached Yes (temporary)
Durable? No No Uncertain
Who controls it? Benchmarks Cloud providers US policymakers
Anthropic controls? No No No

Here, then, is what the moat analysis actually tells us. Model quality is a narrowing lead. Cloud distribution is already breached. Geopolitics is the most durable barrier — but it is not a barrier Anthropic built, controls, or can count on. A shrinking benchmark gap, a distribution channel that lists competitors, and a trade policy you do not own. That is the foundation for a $900 billion valuation.


The Books

Anthropic posted $10.9 billion in Q2 revenue and $559 million in operating profit — its first profitable quarter ever. The confidential IPO filing followed in June 2026, targeting a $900 billion-plus valuation.

The timing is not subtle. Revenue is reported gross, including cloud partner shares. Claude Code drives growth by routing developer token consumption through Anthropic’s API — a genuine distribution engine that works equally well with any model backend once coding parity arrives.

OpenAI reported $5.7 billion in Q1 revenue. Operating margin: negative 122 percent. Projected full-year loss: $14 billion. Cumulative losses through 2028: $44 billion. Nine hundred million weekly users at roughly $2 per month each. They have the audience. They do not have the economics.

Zhipu — traded as z.ai on the Hong Kong exchange — is the cautionary tale. GLM-5.2 is a genuinely competitive model. It is 1–4% behind Claude on coding benchmarks at a fraction of the price. But the business tells a different story: $105 million in 2025 revenue against a $438 million adjusted net loss (¥3.18 billion; total net loss was ¥4.72 billion, approximately $650 million). 2 The stock went from HK$116 at its January IPO to a June peak of HK$2,980 — a 1,280× price-to-sales multiple on a 4% free float — then crashed 70% to HK$905 today. The triggers were a botched HK$31.4 billion placement, the Kimi K3 launch, and lockup expiry. But the underlying problem was simpler: you cannot sustain a trillion-HKD valuation on $100 million in revenue.

The market overreacted on the way up and is likely overreacting on the way down. But the point stands: the gap between $105 million and $45 billion is not explained by better benchmarks. It is explained by enterprise relationships, English-language developer mindshare, and US market access — advantages with nothing to do with model architecture.

DeepSeek does not publish financials. It is not public. It may never be. The company is funded by High-Flyer, a quantitative hedge fund. The model is open weights, MIT license. V4 Pro’s price was cut 75% in May and made permanent. DeepSeek does not need its models to be profitable. The models are a byproduct. The real business is trading. Everything else is infrastructure.


The SPCX Pattern

If you want to know how the Anthropic IPO plays out, you do not need a crystal ball. You just need to look at SpaceX.

SpaceX went public in June 2025 at $135 per share. Three days later it hit $225 — a $3 trillion market cap on a 4.2% free float. Today it trades around $108 to $118, below its IPO price.

The pattern: extreme float scarcity creates a pop — demand overwhelms supply, the price disconnects from valuation. Insiders are locked up, the stock runs on narrative. Then lockups expire — 911.5 million restricted shares in SpaceX’s case, potentially expanding the float nine times. Insiders cash out. Fundamentals reassert.

SpaceX’s fundamentals, by the way: Starlink generates $4.4 billion in operating profit. The AI business loses $6.4 billion. The net was a $4.9 billion loss in 2025, accelerating to $4.3 billion in Q1 2026 alone. Morningstar’s fair value estimate: $63 per share — roughly half the IPO price.

Anthropic is running the same script: turn profitable, file confidentially, float a small percentage, let scarcity drive the valuation. The question is whether the fundamentals — a 5% margin on $10.9 billion in revenue, $100 billion-plus in cloud commitments, training costs that compound with each generation, and competitors closing the gap at 1/50th the price — can support the narrative long enough for the lockups to expire.

The market tends to figure these things out before the insiders do.


Who Is Actually Going to Fund Advanced Research?

At this point a reasonable question arises: if the model business is so bad, why does anyone build models? Who funds the research?

The answer is already visible in the market structure. Look at who is actually doing frontier work and why.

Who Funds It Real Business Model Is A…
DeepSeek / High-Flyer Quantitative trading Cost center (R&D for trading)
Meta Advertising Cost center (commoditize complement)
Google Cloud compute Cost center (sell more VMs/TPUs)
Microsoft Cloud + Enterprise Cost center (sell Azure + Copilot)
Amazon Cloud compute Cost center (sell more Bedrock)
Anthropic API billing THE business
OpenAI Subscriptions + API THE business
Zhipu API billing THE business

The companies winning the model race are not model companies. DeepSeek is funded by a hedge fund that does not need to make money on API calls — it needs better models for trading. If open-sourcing those models collapses Western pricing power, that is a side effect. Meta gives away Llama to protect its advertising business; if models are free, Google cannot use AI as a wedge. Google and Amazon host every model because every inference call is a compute sale. The model is a cost center that enables a profit center elsewhere.

The only companies trying to profit from selling the model itself are Anthropic, OpenAI, and Zhipu. Two are losing billions. The third just recorded its first quarterly profit — right before filing to go public.

Frontier AI research will continue. It will be funded by companies whose real business is something else. The model is infrastructure. Infrastructure is funded by the businesses built on top of it. That is how every technology stack in history has worked.


The ISP Analogy

The closest historical parallel is the internet infrastructure buildout of the late 1990s. A vast amount of capital went into laying fiber. The promise of endless bandwidth was correct. The bet on who would capture the value was wrong.

Internet Era AI Era
Google · Meta · Amazon [Being built right now]
Netflix · Salesforce · Stripe The AI-native companies
ISPs · Fiber · CDNs (Essential. Low margin. No pricing power.) Model APIs · Compute (Essential. Margin compressing fast.)

The ISPs and fiber companies did not go to zero. Some consolidated. Some went bankrupt. Some survived as regulated utilities with single-digit margins and no pricing power. The companies that captured the value — Google, Netflix, Amazon, Salesforce — built on top of the infrastructure. They did not own the pipes. They owned the relationship with the customer, the data, the workflow, the brand.

Model companies are fiber. They are building essential infrastructure that gets cheaper every month. The trillion-dollar outcomes will belong to the companies that figure out how to build on it — not the companies selling tokens.

Fiber at least had physical moats. The switching cost between model APIs is an API key.

And here is the cruel part: you cannot dig up Manhattan to lay competing cable. You can, however, train a competing model if you have GPUs and data. The “moat” in AI is thinner than anything the ISP industry ever had.


Cost Structure Is the Real Disruption

Most of the AI conversation is about quality. Can the model write better code? Reason more deeply? Pass harder benchmarks?

This is the wrong question.

AI is not winning because it is better. It is winning because it is several hundred times cheaper. And cost structure disruption is far more dangerous to incumbents than quality disruption — because the incumbent literally cannot respond.

AI isn’t winning because it’s better. It’s winning because it’s a million times cheaper.

Watch what happens when AI is good enough and costs approach zero:

Task Human cost AI cost (DeepSeek) Multiplier
Review 1,000 legal contracts ~$30K (junior lawyers, 1 week) ~$0.03 1,000,000×
Handle 100K support tickets/month ~$2M/year (50-person team) ~$5K/year 400×
Translate 1 million words ~$150K (professional service) ~$0.50 300,000×
Code review 10K PRs/month ~$750K/year (5 senior engineers) ~$2K/year 375×

These are not edge cases. These are core operational costs across legal, support, localization, and engineering — the cost centers that define professional services, SaaS, and financial services margins.

Now here is why incumbents are structurally unable to capture this:

Incumbent Economics AI-Native Economics
Revenue model Billable hours Outcome-based pricing
Cost base Headcount-heavy Compute + small team
Culture Pyramid structure Lean, automated
Incentive Grow headcount Reduce headcount

If they adopt AI and cut costs 90%: Revenue collapses (hourly billing × fewer hours). The partnership model breaks — no associates, no partners. The culture rebels — the work they trained for disappears.

If they don’t: Someone else does it, charges a tenth of the price. Their best people leave. Their clients follow.

This is Clayton Christensen’s Innovator’s Dilemma, except the disruption is not “cheaper but worse and improving over time.” The improving part already happened. We are at the part where the technology is good enough now and incumbents are structurally paralyzed.

A three-partner law firm with AI can do the work of 200 lawyers, charge a fraction, and have higher margins. A five-person audit firm with AI handles more engagements than a Big Four office. An insurance startup with automated claims processing operates at cost structures that make incumbents’ combined ratios look like a different industry.

Nobody can beat you by being 10% better at contract review. But someone will absolutely destroy you by being a million times cheaper — even if they are slightly worse. Most customers will accept “good enough” in exchange for a price that rounds to zero.

The winning strategy is not “add AI to your business.” It is “rebuild your business around AI’s cost structure.” If you are not willing to cannibalize your own headcount-based revenue model, someone else will do it for you. And they will not give you a warning shot.


What This Means

This analysis has implications for three groups.

If you are building an AI startup: your costs are going down, not up. Do not budget like your model bill will grow with usage. Budget like it will shrink. Architect your system to be model-agnostic — your business logic should survive any specific model being replaced next quarter. Invest in what is defensible: your data, your domain expertise, your customer relationships, your workflow integration. Models are interchangeable. Your understanding of your industry is not.

If you lead a company adopting AI: the killer app is not a chatbot on your product. It is rebuilding a core process so it costs 1/500th of what it costs today. Look at your P&L. Find the biggest line items that are fundamentally headcount-driven. Those are your targets. And ask yourself honestly: if a competitor rebuilt their entire cost structure around AI and could charge a fraction of what you charge, could you respond without destroying your own margins? If the answer is no, you have a strategy problem, not a technology problem.

If you are considering investing in AI companies: model companies at 1,280× revenue are not investments. They are speculative bets on a narrative that is already crumbling. The value in AI does not accrue to the companies that build the models. It accrues to the companies that build on them — the ones using cheap, commoditizing infrastructure to deliver outcomes at cost structures that were impossible two years ago. Those are the companies worth finding.


The AI revolution is real. So is the margin compression. The profit will not go where most people think.

❦ ❦ ❦

Sources & References

Token share & OpenRouter data:

Enterprise adoption:

Cloud marketplace availability: 1

Company financials:

SpaceX / SPCX:

Note on data scope: Token share figures throughout this article are drawn from OpenRouter platform telemetry, which represents approximately 1-2% of global API token traffic and skews toward price-sensitive developers. Token volume share does not equal revenue share — Anthropic captures more than half of OpenRouter spending on approximately 12-13% of token volume. These figures are best understood as a directional signal, not a census of all enterprise AI usage.

Footnotes

  1. DeepSeek-R1 was available on Bedrock Marketplace from January 30, 2025, and as a fully managed, generally available serverless model from March 10, 2025. This article originally stated “since February 2026,” which refers specifically to the V3.2 model addition. The corrected timeline reflects the earliest generally available DeepSeek model on Bedrock. 2

  2. Zhipu reported total net losses of ¥4.72 billion (~$650 million) for 2025. The $438 million figure cited in this article refers to the adjusted net loss of ¥3.18 billion, which excludes certain one-time and non-cash items. Both figures are from Zhipu’s post-IPO financial disclosure as reported by SCMP. 2