Independent business intelligence · Published by Magrofy Inc.New insights every week
Look At BusinessLook At BusinessLook At BusinessBusiness intelligence
Partner

Your Moat Just Got 43% Cheaper to Cross

Explore how falling technology, distribution, and execution costs can weaken competitive moats and make established markets easier for new entrants to attack.

By Editorial TeamAugust 11, 2026
Your Moat Just Got 43% Cheaper to Cross
FOUNDERS & ENTREPRENEURSHIP

Average inference pricing fell from $2.04 to about $1.17 in ten weeks. If any part of your defensibility rested on affording model spend a competitor couldn't, that ended over one summer — and you probably haven't repriced.


Average inference pricing fell from $2.04 to about $1.17 in ten weeks. If any part of your defensibility rested on affording model spend a competitor couldn't, that ended over one summer — and you probably haven't repriced.

LookatBusinessAnalysisFor founders & early-stage CEOsAugust 12, 2026

In late May, average inference pricing sat at $2.04 per million tokens. By late July it was $1.45. In the first week of August it reached $1.16 to $1.18 — a 2026 low, and roughly a 43 percent decline in ten weeks.

Three things drove it at once: aggressive rate cuts from the largest providers, a flagship-class model arriving at roughly half the previous price point, and Chinese open-weight models taking enterprise share and dragging the floor down behind them.

Most founders read that news as good. Their unit economics improved without them doing anything. That reading is correct and incomplete, and the incomplete half is the one that matters.

The barrier to entry fell faster than your costs did

When the cost of a critical input drops by 43 percent, it drops for everyone. Your margin improved. So did your competitor's. So, more importantly, did the economics of a two-person team building the thin version of what you sell.

And here is the asymmetry: you carry the organization you built when it was expensive. The headcount, the infrastructure, the processes, the burn. The new entrant carries none of it. They get the new price and none of the legacy.

A cost advantage in a market with a published price list and a falling curve is not a moat. It is a lead time — and unlike a moat, a lead time is a wasting asset whether or not you use it.

This is worth being precise about, because "we can afford to run this at scale and they can't" was a genuine and reasonable advantage in 2024. Serving a model-heavy product at volume required either capital or unusual efficiency. That was a real barrier. It was also, in retrospect, a barrier priced by someone else's pricing page, which meant it could be removed by an announcement.

What survives a price collapse

The useful exercise for any founder this month is to list the things you believe make you hard to displace, and sort them by whether they have a public price.

Visual 1 — Sorted by whether someone can simply buy it

Claimed advantage

Has a public price?

Survives a 43% input-cost cut?

We can afford the compute at scale

Yes — it is a rate card

No. This one just evaporated.

We use the best available model

Yes, and it changes quarterly

No. Model access is not differentiation.

Our prompts and pipeline are sophisticated

No, but replicable in weeks

Weakly. Buys months, not years.

Proprietary data nobody else has

No

Yes — if it is genuinely proprietary and genuinely useful

We are embedded in the customer's workflow

No

Yes. Switching cost is paid by the customer, not by you.

Distribution — we reach buyers others cannot

No

Yes, and it appreciates rather than decays

Regulatory position, certification, approvals

No — priced in time, not money

Yes, and time is the one input nobody can discount

How to read it: The top two rows are what a great many 2024 and 2025 pitch decks described as defensibility. The bottom four are what actually holds. The question is not whether you have any of the bottom four — it is whether you were building them while the model spend was doing the work of looking like a moat.

That last framing is the one worth sitting with. Expensive inputs are flattering. They make a product feel substantial and a competitor feel far away. They can also conceal, for a couple of years, the fact that nothing durable is being accumulated underneath — and the concealment ends the moment the input gets cheap.

Two traps in the obvious response

The first is spending the savings. The natural reaction to cheaper inference is to use more of it — richer context, more calls per interaction, more agentic steps. Each is individually justifiable, and collectively they are why cheaper inputs so often produce larger bills. Unit costs falling while total spend rises is one of the oldest patterns in economics, and it is playing out here in real time: across the market, inference spending has now overtaken training spending.

The complication is that most companies cannot see it happening. Recent research found that fewer than a third of organizations have accurate visibility into their AI software spend — against roughly three-quarters for on-premises hardware. If you cannot see the line, you cannot notice it doubling.

The second is assuming your pricing holds. If your price to customers was set when your costs were 43 percent higher, you now have a margin windfall — and so does everyone else, which means someone will use it to undercut you. Better to decide deliberately whether that windfall becomes margin, becomes price, or becomes investment in something that does not have a rate card, than to discover the decision has been made for you by a competitor's pricing page.

What to do this month

Recompute gross margin at today's prices, and then ask what your price should be. Not what it is — what it should be, given the new cost base and what a well-funded entrant could charge. If the answer is uncomfortable, better to know now.

Name the part of your defensibility that has a public price. Be honest about the proportion. If most of it does, you are running a lead-time business, and lead time should be spent buying something durable rather than banked.

Instrument the spend before you expand usage. Cost per active customer per month, tracked weekly. It takes an afternoon while the volume is small and becomes very hard once it is not.

Write down what must be true in 2028. Assume inference is another 50 percent cheaper and that model capability is broadly commoditized at the level you use. What is left that makes a customer choose you? That paragraph is your actual strategy, and if it is difficult to write, that difficulty is the finding.

Reconsider anything you rejected as too expensive in 2025. The curve moved. Ideas that were uneconomic eighteen months ago may not be now — and this is the genuinely good news in the story, provided you go looking for it rather than waiting for someone else to.

Every founder has read that cost advantages are weak moats. It is one of those things everyone agrees with in the abstract and quietly exempts themselves from, because their cost advantage feels structural rather than circumstantial. Ten weeks in one summer settled the question. The advantage was circumstantial. What you built underneath it, while it lasted, is the only part that was ever going to matter.


Sources and method. A LookatBusiness original. Inference pricing figures — an average of $2.04 per million tokens on May 31, 2026, $1.45 in late July, and $1.16 to $1.18 on August 6–8, 2026 — are from Jefferies research using Silicon Data, as reported by the South China Morning Post, August 10, 2026, which attributes the decline to provider rate cuts, a flagship-class model released at roughly half the prior price, and Chinese open-weight models taking enterprise share. Inference spending overtaking training spending within AI-optimized infrastructure per Gartner, August 10, 2026 — a market forecast, not a measurement. Visibility into AI software spend (fewer than one third of organizations with accurate visibility, against roughly three-quarters for on-premises hardware) from Flexera's 2026 State of ITAM research, published July 2026 — vendor research with a disclosed respondent base. Prices in this market move quickly; figures are dated deliberately and should be re-checked before being relied on.