The wrapper argument is incomplete
A user does not buy a database; the user buys accounting, logistics, publishing, customer support, design, or analysis. The product layer defines the job, gathers context, creates interfaces, enforces permissions, handles exceptions, and connects output to action. AI does not change that logic.
The weak product wraps a model response in branding and charges for access. The strong product wraps a complete job: it knows what starts the work, which information is authoritative, what quality looks like, which actions are allowed, how humans review exceptions, and what outcome the customer measures.
This is why a “wrapper” can become a durable business and a technically sophisticated model product can fail. The question is not how much original machine learning the company owns. It is how much of the customer’s valuable system it improves and how difficult that improvement is to reproduce.
Prompt quality matters, but it is one component inside context, tools, permissions, evaluation, interface, distribution, support, and accountability.
Model capability is commoditizing
The Stanford AI Index documented a more than 280-fold reduction between late 2022 and late 2024 in the cost of querying a model that reached a GPT-3.5-equivalent MMLU score. The report also found smaller models reaching capability thresholds that previously required far larger systems. Prices and architectures have continued changing since then.
This does not mean frontier models are identical or that model choice is irrelevant. It means a product cannot assume permanent scarcity. A feature that requires an expensive model today may run on a smaller, cheaper, local, or open model later. A provider may add the feature directly. Competitors may route across models.
Model dependency creates three strategic risks:
- Capability risk: the provider improves, deprecates, or changes behavior faster than the product can adapt.
- Economic risk: price, rate limits, context costs, or latency make the current workflow unprofitable.
- Platform risk: the provider launches a native product that captures the same user relationship.
The defense is not necessarily training a foundation model. It is building a layer that becomes more useful as models improve because the workflow, customers, data, and trust remain yours.
Moat 1: Workflow ownership
Workflow ownership means the product solves a complete, recurring job rather than generating an isolated artifact.
A ticket arrives, a document changes, a lead qualifies, a file is uploaded, or a threshold is crossed.
Retrieve approved data, policies, history, examples, and constraints.
Separate deterministic rules, model inference, and human discretion.
Create, route, update, schedule, export, or request approval.
Detect uncertainty, escalate, ask for missing context, and preserve a fallback.
Track acceptance, resolution, revenue, defects, cycle time, or another business signal.
A generic writing tool competes with every model. A product that understands a regulated claim-review process, retrieves the right evidence, enforces required fields, produces an auditable packet, and routes exceptions competes on a system.
Workflow depth also creates retention. The customer has invested in templates, integrations, approvals, history, and team behavior. That switching cost is legitimate when it reflects accumulated value—and abusive when the product blocks export or hides data.
Moat 2: Distribution
A technically strong product without a repeatable path to users is a project. AI increases this pressure because competitors can build similar demos quickly.
Distribution is more than advertising. It includes:
- ownership of a trusted audience or community;
- presence inside a system where the workflow already occurs;
- partnerships, marketplaces, or channels with aligned incentives;
- content that answers the problem before the buyer selects a tool;
- team adoption that spreads through shared artifacts and collaboration;
- sales and onboarding suited to the buyer’s risk and complexity.
The strongest distribution is product-shaped. A useful output travels: a report, link, file, workflow, template, or collaboration invites another user. The product can also sit at the moment of need—for example, inside a help desk, code repository, browser, document system, or vertical platform.
Paid acquisition can start growth, but it rarely protects a product whose value is easy to evaluate and easy to replace. When model interfaces converge, trusted problem ownership becomes the brand.
Moat 3: Feedback and evaluation data
“Proprietary data” is often used vaguely. A pile of customer documents is not automatically a moat, and collecting more sensitive material can create liability. The valuable data is often structured feedback tied to outcomes:
- which outputs were accepted, edited, rejected, or escalated;
- which errors were severe and why;
- which context resolved the task;
- which user, segment, or scenario changed performance;
- what happened downstream after the output was used;
- how the workflow changed over time.
This data improves prompts, retrieval, routing, user experience, and evaluation. It can also reveal that the model is not the bottleneck.
An evaluation set is a product asset because it defines quality in the customer’s domain. When a provider changes a model, the company can test whether the product improved or regressed. Without that evidence, each model update becomes a live experiment on customers.
Do not call data a moat if customers cannot understand or control how it is used. Minimize collection, separate product telemetry from training consent, protect sensitive examples, and prefer outcome labels over indiscriminate content hoarding.
Moat 4: Trust and operational reliability
Trust is not a marketing adjective. It is the accumulated evidence that the product behaves predictably, protects data, admits limits, supports recovery, and remains accountable when something goes wrong.
NIST’s Generative AI Profile identifies risks such as confabulation, privacy, information integrity, security, bias, and overreliance. A production product needs controls for the subset that matters to its use case:
- grounding and source links where factual claims matter;
- validation rules for structured output;
- permissions and approval gates for actions;
- clear retention and training terms;
- logs, versioning, and incident response;
- fallbacks and export;
- human support that understands the workflow;
- honest capability boundaries.
Trust compounds slowly and can collapse quickly. That makes it a durable advantage when earned. A provider can copy a button; it cannot instantly copy years of reliable outcomes, domain support, security review, and customer confidence.
The common failure modes
The product generates impressive output but does not own the trigger, context, review, action, or outcome.
The product claims to serve every team, so it fits none of their systems or quality standards deeply.
Users discover the product only through expensive ads while alternatives multiply.
The company sees prompts and outputs but does not learn which results were correct or valuable.
Fluent output requires so much verification that the workflow becomes slower or riskier.
A model price, policy, or native feature change removes the economics or differentiation.
The product makes consequential claims without evidence, audit, security, or recovery.
Customers stay because data and workflows are trapped, not because the product improves.
Procurement approves the product, but workers bypass it because it adds friction or misses the real job.
These failures interact. Weak workflow fit lowers usage, which produces little outcome data, which prevents improvement, which makes distribution more expensive, which pushes the company toward exaggerated claims.
Unit economics after review
AI product economics should be calculated per accepted outcome, not per generated token or active seat.
| Cost | Why it matters | Product response |
|---|---|---|
| Inference | Long context, retries, tools, and reasoning can multiply base token cost | Route tasks, cache, use smaller models, constrain context |
| Review | Human verification can exceed generation time | Improve grounding, validation, interface, and confidence signals |
| Error | One severe failure can dominate thousands of cheap successes | Weight severity, add approvals, narrow automation |
| Integration | Connectors, permissions, maintenance, and support are ongoing | Own common integrations and standardize workflows |
| Acquisition | Easy-to-copy features increase paid-marketing pressure | Build channel, product-led spread, or domain authority |
| Churn | Users leave when novelty fades or a platform bundles the feature | Deliver recurring workflow value and accumulated context |
METR’s finding that experienced developers took longer with AI in a controlled study is a reminder that perceived speed is not enough. A product must measure completed work, not the feeling of acceleration.
A product survival scorecard
Does the product own a repeated job from trigger to outcome, including exceptions?
Is there a repeatable channel, embedded position, audience, or product-led loop?
Does use generate lawful, outcome-linked evidence that improves the product?
Can the company prove data handling, quality, control, auditability, and recovery?
Can the product route, switch, or continue if an upstream provider changes?
Below 50: feature risk. 50–70: useful product but exposed. Above 70: compounding system—assuming the underlying market is real.
The score does not replace customer evidence. It forces a team to name what will remain valuable after the next model release.
Why focused utilities can still win
Not every durable product needs proprietary training data or a large enterprise platform. A focused utility can win through clarity, speed, privacy, search distribution, and excellent execution.
Jivaro’s browser apps are examples of a deliberately narrow strategy: compare documents, resize images, test regex, preview metadata, edit Markdown, split text, or run a frontend experiment. Many of these tasks do not need generative AI. The product advantage is that the user can open a URL, complete the job locally, and leave with a portable result.
That approach has limits—utilities can be copied, and many users are infrequent—but it avoids a common AI failure: adding probabilistic complexity to a deterministic problem.
Use AI where ambiguity, language, or inference creates real value. Use deterministic browser processing where the job is transformation, comparison, validation, or export. The product should be judged by the solved workflow, not by how much AI it contains.
What the market rewards next
Cheaper models make raw capability less scarce; products that own execution, context, and outcomes capture more value.
Customers will expect visible tests, monitoring, model/version history, and evidence that updates do not regress critical work.
Products embedded in a role or industry can accumulate deeper context, integration, and trust.
Many applications will select among frontier, small, local, or specialized models based on task, cost, latency, and risk.
Security, provenance, permissions, recovery, and support become buying criteria as agents take real actions.
Features without workflow or distribution moats will be bundled by model providers, operating systems, or established SaaS vendors.
Frequently asked questions
No. Wrapping a model is normal product development. The risk is adding no durable workflow, distribution, feedback, or trust layer above the model.
Not always. Proprietary outcome feedback and evaluation can be more valuable than raw content. Some products win through distribution, workflow depth, privacy, or execution without training on customer data.
For most application companies, workflow ownership is the foundation because it creates integration, retention, feedback, and measurable outcomes. Distribution and trust often determine whether that workflow becomes a business.
Demos use ideal inputs and ignore permissions, integration, latency, review, edge cases, severe errors, support, and downstream outcomes.
It can copy features and interfaces. It is harder to copy deep customer workflows, integrations, domain feedback, support, distribution relationships, and accumulated trust.
No. Add AI when inference improves the user’s job enough to justify variance, cost, and governance. Deterministic tools are often better for conversion, validation, calculation, comparison, and export.
Sources and references
- Stanford HAI: 2025 AI Index—Research and DevelopmentStanford HAI · reference
- NIST: Generative AI ProfileNIST · reference
- METR: Measuring AI impact on experienced developer productivityMETR · reference
- U.S. Census Bureau: The Microstructure of AI DiffusionU.S. Census Bureau · reference
- OECD: AI adoption by firmsOECD · reference
- NBER: Generative AI at WorkNational Bureau of Economic Research · reference
- Menlo Ventures: 2025 State of Generative AI in the EnterpriseMenlo Ventures · reference

