← All articles

Strategy

Your engineers just made your biggest pricing decision

Two AI products can charge the same price and keep very different shares of it. The difference gets decided in the architecture — and nobody calls it a pricing decision.

In October 2023, the Wall Street Journal reported that GitHub Copilot — which charged individuals $10 a month — was losing an average of $20 a month per user in the first months of that year. Some users cost Microsoft as much as $80 a month.

That story became the founding myth of AI unit economics: the more people love it, the more you lose. It's still a useful warning. But it hides the more interesting part — how much of your margin is decided by choices nobody files under "pricing."

The margin gap is real

The data has caught up with the anecdote.

  • Bessemer, State of AI 2025 — the fastest-growing AI startups, which it calls "Supernovas," average about 25% gross margin. Steadier "Shooting Stars" average about 60%.
  • ICONIQ, State of AI 2026 — AI product gross margins rising from 45% in 2025 to a projected 53% in 2026 and 59% in 2027.
  • Duolingo, Q1 2025 — gross margin of 71.1%, down from 73.0% a year earlier, attributed to "increased generative AI costs related to the expansion of our Duolingo Max tier."
  • Microsoft, quarter ending December 2025 — Microsoft Cloud gross margin fell to 67%, "driven by continued investments in AI infrastructure and growing AI product usage."

For years, SaaS didn't have to think about this. One more user cost almost nothing to serve. I wrote about what breaking that assumption does to your pricing model in Pricing is the feature you forgot to ship. This piece is the other side of the same coin: the cost line — and who actually controls it.

The levers are on your provider's price list

Model providers sell the same tokens at very different prices depending on how you ask for them. From Anthropic's pricing page:

  • Prompt caching — store a repeated block of context once, then reuse it. Writing to the cache costs 1.25x the normal input price for a 5-minute cache, or 2x for a 1-hour cache. Every read after that costs 0.1x — and on Anthropic's newest models, 0.05x or less.
  • Batch processing — anything that can wait gets "a 50% discount on both input and output tokens." And the two discounts "can be combined."

OpenAI's pricing page shows the same shape: cached input at a tenth of the normal price on its current GPT-5.6 models, and "Save 50% on inputs and outputs with the Batch API."

Here's what that does to one workload. Say your product sends the same large context — a codebase, a contract, a long set of instructions — 100 times within the cache window.

Cost, in units of one normal read
No caching100
Caching: 1 write at 1.25 + 99 reads at 0.111.15

Same work. Same output. About 89% less on that input.

Agent products live inside this pattern. The model doesn't remember the conversation between calls, so every turn sends the whole history again. That's repeated context — the exact thing caching discounts.

The decision nobody puts on the roadmap

Once engineering ships caching, the cost of serving a customer drops. And a question appears that nobody scheduled: who gets the savings?

Keep the arithmetic simple and assume the whole bill is input. You charge a customer $120 for a workload that costs you $100 at standard rates — a 17% margin. Now 90% of that input is served from cache. The provider bill drops to about $21.50: $9 for the cached reads, $12.50 for writing the rest to cache. Charge the same $120, and your margin is about 82%.

You have two honest options.

Pass the savings downKeep the savings
What you doLower the price, or charge cost plus a fixed markupCharge the standard rate; keep the efficiency
What the customer seesA bill that shrinks over timeNo change
Your marginThin, driven by volumeWide
When it fitsYou're fighting for share; your buyers are developers who benchmark unit costsYou're differentiated; your buyers pay for outcomes, not tokens

Neither is wrong. The mistake is not choosing — and most teams don't, because something else chooses for them: the pricing unit.

Charge per token at cost plus a markup, and the savings flow to the customer automatically. Charge per seat, per task, or per outcome, and you keep them. Nobody sat in a room and decided that. It fell out of the unit.

Your pricing unit decides who keeps your engineering savings — so pick the unit knowing that.

Two costs that move the other way

Margins don't only improve on their own. Two things push in the other direction, and they rarely show up in a pricing review.

  • New tokenizers. When Anthropic launched Claude Opus 4.7, it noted that "the same input can map to more tokens—roughly 1.0–1.35× depending on the content type." Same text, same price per token, bigger bill. Every model upgrade needs a cost-per-task benchmark, not just a quality check.
  • Tool fees. Agents pay for their tools on top of tokens. Anthropic's web search costs $10 per 1,000 searches, plus the tokens it pulls in. Code execution runs $0.05 per container-hour once the 1,550 free monthly hours are used up. Thousands of agents searching all day is a line item nobody modeled.

What to bring to your next pricing review

  • Cache hit rate per product surface — is repeated context actually being reused?
  • Workloads that could run in batch and don't — reports, indexing, anything nobody waits for.
  • Gross margin per feature, not just per account — know which features earn and which leak.
  • A cost-per-task benchmark after every model change.
  • A written decision — pass the savings down or keep them, and which pricing unit enforces it.

Your next margin gain might not come from a price increase. It might already be sitting on your provider's price list, waiting for someone in product to claim it.

Who on your team decided what happens to the savings from caching — and did anyone write it down?

Free companion

The Credit Engine — Infrastructure Playbook

The infrastructure behind credits-based pricing in a single PDF — the ledger, metering, and enforcement patterns that decide whether your margins survive once engineering ships the savings.

Get the playbook →

This is one piece of a longer framework I teach in Chapter 5 of Product Strategy in the AI Era — including IQ tiering: gating your most expensive models behind higher tiers, so a power user on a cheap plan can't run up your bill.

Sources

  • AI Business, "GitHub Copilot loses $20 a month per user" (October 11, 2023): aibusiness.com
  • TechRadar, "Microsoft is reportedly losing huge amounts of money on GitHub Copilot" (October 10, 2023): techradar.com
  • Bessemer Venture Partners, "The State of AI 2025" (August 13, 2025): bvp.com
  • ICONIQ, "State of AI 2026": iconiq.com
  • Duolingo Q1 2025 shareholder letter (SEC filing): sec.gov
  • Microsoft FY26 Q2 performance: microsoft.com
  • Anthropic, Claude API pricing: claude.com
  • Anthropic, "Introducing Claude Opus 4.7" (April 16, 2026): anthropic.com
  • OpenAI, API pricing: openai.com

Newsletter

The Strategic AI Corner

One idea on product strategy in the AI era — weekly. No spam, unsubscribe anytime.

Free. Unsubscribe anytime.

More articles

📊
Strategy

Figma gave customers three months to learn its AI credits. One still ended up at 209% of the limit.

Measuring usage before you charge for it is the right move. Figma did it by the book — and its forum still shows exactly where warm-up periods break.

Read →
📋
Strategy

Your customer's procurement team now has a checklist for your AI pricing

Buyers have spent two years absorbing AI price increases. Now they're organizing. Here's what they'll ask — and how to design pricing that passes before they do.

Read →
🤝
Strategy

Your usage pricing won't fail with customers. It'll fail in the comp plan.

With seats, the deal ends at signature. With usage pricing, it starts there — and most sales teams are still paid for the old ending.

Read →
🧾
Craft

Stripe didn't buy a billing company. It bought the line between product and business.

Credits look like a finance detail. Each one carries a dozen product decisions — and in 2025, the companies that got them wrong found out in public.

Read →
🧭
Strategy

Anthropic pricing and model — a lesson for all respected SaaS

Anthropic's top plan unlocks four things, and not one of them is a feature. A masterclass in AI-era segmentation — and what it means for your own pricing.

Read →
📝
Craft

Your PRD now has three readers — and one of them is an AI

PRDs didn't die. They got a new audience — and old habits don't survive it.

Read →
💰
Strategy

Pricing is the feature you forgot to ship

Why pricing decisions belong in the product roadmap, not in finance.

Read →
🎯
Positioning

Every AI feature you add is a positioning decision you didn't make

The hidden cost of adding AI features without a positioning frame.

Read →