Draft for Tamir's review. Not published.
Measure cost and quality by request type
Break model spend down by request type and customer segment. Cost usually concentrates in a few request types, so the fix is often local: smaller models for routine steps, caching, shorter contexts or limits on the heaviest use.
When the CFO wants cheaper models and the CTO fears quality loss, run the cheaper models on the evaluation set and a controlled slice of traffic. Move only the request types that hold. Exhaust prompts, retrieval and evaluation before fine-tuning, and fine-tune only if you can own the retraining cycle.
Keep the dependence one you can undo
Keep tool definitions, evaluation, state handling, prompts and logs provider-neutral, and wrap any framework behind your own interfaces. Build that abstraction while the switching cost is still small. Run your evaluation suite on a second provider and record the quality gap.
Credits, distribution deals and discounted compute all trade money now for dependence later. Accept exclusivity only for a short term that ends with the credits, with the right to evaluate others. Commit compute against contracted demand, and compare self-hosting with hosted inference under the same service expectations.
Depending on your seat
If you're on the board or an investor, don't pick the architecture. Ask for cost per customer and per request, the share going to model providers, and the margin if provider prices change. Require a runway forecast that shows usage and supplier-cost variation.
If you're the CEO, ask whether each account carries manual setup and correction work the sales model left out. If a provider withdraws a capability you depend on, recheck the product's value before choosing a replacement. A narrower product may serve customers better than a full rebuild.
What to check before you decide
- Break model spend down by request type and customer segment, and find where most of it sits.
- Run cheaper models on your evaluation set and a controlled slice of traffic, by request type.
- List every provider-specific feature you use and mark which ones have an equivalent elsewhere.
- Run your evaluation suite on a second provider and record the quality gap.
- Read provider terms on price changes, deprecations, exclusivity and data use, with the notice periods.
- Model the margin after credits end and if provider prices change.
- Separate contracted demand from hoped-for demand before committing compute or buying hardware.
Questions people ask
Evaluating an AI company with strong growth and opaque compute costs, how do I assess the unit economics before investing?
Ask for cost per customer and per request by segment, the share that goes to model providers, the architecture's levers to reduce it, and the dependence on provider pricing. Compare with what customers pay and with the gross margin at scale. It depends on how much of the product's cost is model calls and on whether the company controls that cost.
Management asked the board to choose between building our own models and building on providers, how should the board decide?
The board should not pick the architecture; it should require management to show the decision's consequences: cost, talent, time to product, dependence and what customers gain, under each path, and approve the path with the evidence. Ask what would make management switch later. It depends on the company's data, capital and customers' needs.
A model provider offers credits and co marketing for exclusivity, should the startup board approve?
Approve only if exclusivity has a short term, ends when credits end, excludes the right to evaluate alternatives, and the product's architecture keeps a switch possible. Credits that buy a permanent dependence cost more than they give. It depends on the credits' size against the runway and on how much of the product's quality rests on that provider.
Should the board extend runway forecasts that ignore inference cost volatility?
Require a forecast that exposes plausible usage and supplier-cost variation. The decision depends on workload mix, controllable limits, customer commitments, and finance-owned interpretation of the resulting cash needs.
A big platform offers us distribution if we build on their tools and share revenue, is the dependence worth it?
Take it if the integration can be kept thin, the customer relationship stays yours and the terms have protection against unilateral changes. Platforms change rules; your product should survive that. It depends on how much of your product would need to move onto the platform's tools and on what share of your growth would come through it.
CFO wants cheaper AI models for margin and CTO says quality will drop, how do I decide?
Ask for the test: run the cheaper models on the product's evaluation set and on a slice of real traffic, and measure quality and cost together. Usually some request types can move and some cannot. It depends on whether the company has an evaluation set that reflects what customers notice.
Why do impressive AI demos turn into unprofitable customer accounts?
Rebuild the account economics around actual service obligations. Expansion depends on repeatable onboarding, exception work, usage costs, and whether pricing and product scope match the delivered service.
Should we rebuild our AI product after its main provider withdraws support?
Recheck the product's value before choosing a technical replacement. The decision depends on customer obligations, substitute capability, and whether the revised economics justify continuing the affected offer.
Should our AI startup commit to compute capacity before customer demand is firm?
Compare the commitment with dependable demand and the cost of preserving flexibility. It depends on utilisation, cash exposure, service needs, and the actual cancellation, substitution, or transfer terms.
Our AI product depends on one model provider, should we build multi-model support before enterprise customers ask?
Build the abstraction when the switching cost is still small, which is usually before the second provider is needed. Keep prompts, evaluations and logs provider-neutral so a switch is a test run rather than a rewrite. It depends on how much of your product's quality comes from provider-specific features you cannot replace.
Should we fine tune a model on our domain data or keep improving prompts and retrieval?
Exhaust prompts, retrieval and evaluation first, because they are cheaper to change and show you where the real quality gap is. Fine-tune when the gap is in style, format or consistent behavior rather than in knowledge, and when you can own the retraining cycle. It depends on your evaluation results and on whether your data changes often.
Our AI product costs more per user in model calls than users pay, how do we fix unit economics without losing quality?
Measure cost per task and find the few request types that drive most of the spend, because the fix is usually local. Options include smaller models for routine steps, caching, shorter contexts and limits on the heaviest use. It depends on where your cost concentrates and on which quality your users actually notice.
Should we build our own agent orchestration layer or adopt an open framework that keeps changing?
Keep the parts that make your product different, such as tool definitions, evaluation and state handling, in your own code, and use a framework only where it saves real work and can be swapped. Pin versions and wrap the framework behind your own interfaces. It depends on how much orchestration your product needs and on how much engineering time you can spend on infrastructure.
Our AI product failed on provider rate limits during a customer launch, how do we design resilience without overbuilding?
Start with the cheap controls: know your limits, queue and degrade gracefully, and warn customers before big launches. Add a second provider only for the requests where failure is unacceptable, with quality tested in advance. It depends on which requests must never fail and on what limits the provider will commit to in writing.
When is self-hosting AI models justified for a startup with unpredictable demand?
Compare complete operating requirements across ordinary and peak demand. The choice depends on utilisation, latency, model fit, staffing responsibilities, and the cost of handling overflow.
How I can help with this decision
- Ask or talk (Free)
- I give my view on where cost usually concentrates in products like yours and how much provider dependence you can carry. I tell you which single test shows the real switching cost.
- Review (Pay if it was worth it)
- I write an independent assessment of your unit economics, architecture options and provider dependence, for management, the board or an investor. I recommend which changes to make, in what order, and what to defer.
- Retain (When it makes sense)
- I stay close through the changes to review cost and quality results, and design choices that recreate the dependence in a new place.