On platform dependency: when ‘we use OpenAI’ becomes a deal-killer.
Every AI startup runs on someone else's inference. The question isn't whether — it's how exposed you are when prices triple, capacity gets rationed, or the platform launches a competing feature.
01The premise
“We use OpenAI.” “We use Anthropic.” “We're fine-tuning Llama for the regulated cohort.” Every AI deck contains one of these lines. The line is fine. The question is whether the founder has a real answer to the three risks the dependency introduces, or whether they treat the line itself as the answer. In our pipeline of decks since January, roughly one in four AI startups treats it as the answer. Some of those companies are extraordinary. Most of them are renting a moat they have not yet built.
The pattern matters because platform dependency is asymmetric. The provider can move prices, ration capacity, or ship a competing feature with sixty days' notice. The startup cannot meaningfully respond inside sixty days. A defensible diligence frame for AI investing is to assume the provider will, at some point, do all three — and ask the founder what they've already built to absorb it.
02Risk one — price
Inference costs move. Token prices on frontier models decreased roughly 10× over eighteen months, which sounds like permanent good news until you remember the direction can reverse. A reasoning-model API that ships at $15 per million input tokens can be re-priced at $30 per million with a single update — and several have. The major providers have all moved prices both directions in the past eighteen months. Anyone modelling unit economics on a flat assumption is modelling against the wrong distribution.
The diligence question is specific: at the founder's current gross margin, what increase in input-token price wipes out the margin? If the answer is 20%, the business does not survive a single re-pricing cycle and the founder hasn't built the buffer. If the answer is 200%, the business is robust and either the margin is already high or the inference cost is a small fraction of total COGS.
The founder should know this number cold. Most do not. The ones who do almost always have a fine-tune in their stack — they've already moved the highest-volume routine traffic off the frontier model onto a smaller, cheaper, in-house model, and they can quantify the cost. The ones who do not are running all traffic through the most expensive endpoint and treating the provider's pricing as a constant. The first kind survives a re-pricing. The second kind discovers the term sheet they signed assumed a margin profile that no longer exists.
03Risk two — capacity
Rate limits, tier-cap policies, and account-level quotas are not disclosed in advance and are not, in any meaningful sense, contractual. A company building on a frontier model with a 10K RPM allocation today may discover the allocation halved next quarter, with no recourse other than a higher commitment to the provider — which the provider may not, in fact, offer. This is not a theoretical risk. It has happened to multiple companies in our pipeline in the past nine months, including two we eventually backed.
The diligence signal is whether the founder has a fallback inference path. There are two kinds of answers, and they are not equivalent:
- The good answer is specific. “We route to Anthropic for safety-flagged queries because their content policy is more permissive for our use case. We have a Llama-3.1-70B fine-tune for high-volume routine traffic that handles 60% of our calls at one-fifteenth the cost. Our evaluation harness benchmarks all three providers nightly. We could shift 80% of traffic off any single provider in seven days.”
- The bad answer is generic. “We could swap providers if we had to. The APIs are pretty similar.” This is functionally a no. Swapping providers requires not just code paths but evaluation harnesses, fine-tunes, customer agreements that did not specify a provider, and re-validation across the entire prompt corpus. Companies that can do it in thirty days have built for it from day one. Companies that cannot are exposed and will remain exposed until they re-architect.
The asymmetry: the founder is betting that the provider will not constrain capacity in a way that breaks the product. The provider is making decisions about capacity at the aggregate level, balancing thousands of customers. The founder is not in the room when those decisions are made.
04Risk three — the competing feature
This is the deal-killer. Inference providers are not neutral utilities. They are companies with product roadmaps. The application layer is, increasingly, where they want to play. A startup that defends itself purely on prompt engineering, UX polish, or a clever workflow is defending against a competitor with three things: access to the underlying model the startup is built on, knowledge of the startup's usage patterns, and a pricing power the startup cannot match.
The pattern is repeating. A startup builds a thin wrapper on a frontier model with proprietary prompts and a clean UI. The model provider ships a comparable feature six months later as a free add-on, sometimes integrated directly into the chat interface. The startup discovers their differentiation was the wrapper itself, and the wrapper is now commoditised. The startup's response is usually to add features the provider hasn't shipped — but the provider can also ship those, and frequently does, on a faster cadence than the startup can.
How to read the risk on the deck:
- If the startup's core value is the prompt engineering, the platform owns the moat. The startup is doing the platform's prompt-engineering R&D for free. This is a feature, not a company.
- If the core value is proprietary data, distribution, or workflow integration the platform cannot replicate without acquiring a customer base, the moat is real. Distribution into a regulated vertical, a customer dataset the platform can't access, or a workflow the platform doesn't care to build — these are defensible.
- If the founder cannot articulate why the platform won't build this themselves in twelve months, you have your answer. The right follow-up is to ask what the founder would build if they were the platform's product manager assigned to the category. If the answer matches the founder's actual roadmap, the moat is the customer relationship, not the product.
05What good looks like
A defensible AI startup has at least two of the following five characteristics. The strongest companies have three or four:
- Multi-provider routing with measured per-task quality benchmarks. Not aspirational — actively deployed, with nightly evaluation against the company's real traffic distribution.
- Proprietary fine-tuned models for high-volume use cases. The fine-tune does not have to be exotic. A Llama-3.1 or Mixtral derivative trained on six months of the company's own logged traffic, hosted on the company's own infrastructure, with a measurable cost-per-call advantage over the frontier alternative.
- A data advantage the platform doesn't have. Customer-specific data, regulated data, multi-party data — anything the platform cannot acquire by training on the public web. The data must be both proprietary and durable; a one-time data dump is not a moat.
- Workflow integration that's hard to replicate at the application layer. Embedded into a system of record, an EHR, a CRM, a board portal — the integration cost is the moat.
- Distribution into a vertical the platform isn't pursuing. Energy, healthcare, defence, regulated finance — verticals where the platform's general-purpose product runs into compliance friction the startup has already absorbed.
A startup with none of these is renting a moat. They may run profitably for several years, but they are not building durable equity. The right question for the IC is not whether they have a moat today — it is whether they are converting their current dependency into something the platform cannot replicate.
06The kill flag
We mark platform dependency as a kill flag in our dossiers when two conditions hold simultaneously: the founder cannot describe a thirty-day mitigation plan, AND the core product is inference-dependent in a single-provider configuration. This combination appears in roughly one in four AI decks we analyse. It is not always fatal — but it is always a term-sheet condition. The conversation we recommend is to ask the founder for the mitigation plan up front, before pricing the round. The answer determines what kind of company you're actually backing.
The thirty-day frame matters because re-architecting for multi-provider takes between two weeks and three months of focused engineering effort, depending on the company's current stack. A founder who has not done this work pre-Series A is signing up to do it under pressure, in production, while also building their next quarter's features. That is the kind of structural risk the round is supposed to price in, and rarely does.
Want the rebuild on your TAM slide?
The four-cut framework — and every other pattern in the research archive — is baked into every Early Capital dossier. Upload a deck and you get the reconciliation back in roughly fifteen minutes: claimed number, rebuilt number, the gap, and where it came from.
Concentration: when one customer is half the revenue.
A single design-partner customer can mask everything wrong with the GTM motion. Here's how we test for healthy vs. unhealthy concentration — and the three remediation paths IC tolerates.
Read note →The TAM slide is almost always wrong.
Top-down market sizing is how decks justify ambition. Bottom-up is how investors check the math. The gap between the two is usually where the real conversation happens.
Read note →Three questions every angel should ask, but rarely does.
Most angel diligence calls run forty-five minutes and never get past traction. These three questions surface more than the other forty-five combined.
Read note →