The adoption trap
AI tools create unusually persuasive demonstrations because language and images make capability look general. A strong output in one prompt can feel like proof that the product understands the entire job. Research has repeatedly shown why that inference is unsafe.
In the Generative AI at Work field study, a conversational assistant increased customer-support productivity by 14% on average, with much larger gains for novice and lower-skilled workers. That is meaningful evidence—but it is evidence about a defined support workflow with integrated context, not a universal productivity law.
The “jagged technological frontier” experiment found substantial speed and quality gains on tasks within the model’s capability boundary, while performance could deteriorate when participants trusted the tool on tasks outside it. In a different setting, METR’s randomized study of experienced open-source developers found that early-2025 AI tools made participants slower on their own repositories even though the developers expected the tools to help.
The conclusion is not that AI works or does not work. It is that value is task-specific, user-specific, and process-specific. A credible evaluation must reproduce the environment in which the tool will operate.
Define the job, baseline, users, inputs, acceptable error rate, review process, and output standard before comparing vendors. Otherwise the pilot measures how impressive the interface feels.
Start with the workflow
Begin with a narrow process that has a stable enough baseline to measure. “Use AI in marketing” is not a workflow. “Draft the first version of product-description variants from approved product data, then route them through brand and factual review” is.
What event starts the work: a ticket, document, call, lead, invoice, or request?
Which data, instructions, policies, examples, and tools are required?
What judgment or transformation must occur, and which parts are rules versus discretion?
What output must be produced, in which format, with which evidence or provenance?
Who approves, corrects, signs, publishes, or accepts responsibility for the result?
Measure the current process before adding the tool. Useful baselines include cycle time, labor minutes, first-pass acceptance, defect rate, escalation rate, customer outcome, and the cost of a serious error. If the work has no baseline and no quality definition, the pilot cannot distinguish improvement from novelty.
Prioritize workflows with high repetition, accessible context, reviewable outputs, and reversible errors. Avoid beginning with decisions that are rare, high-stakes, poorly documented, or difficult to verify. The fastest path to failed adoption is to choose a politically visible use case whose outputs cannot be measured.
Map data and failure risk
NIST’s AI Risk Management Framework is useful because it treats AI risk as a lifecycle and organizational problem rather than a one-time security questionnaire. Its companion Generative AI Profile adds risks specific to generative systems, including confabulation, privacy, information integrity, security, bias, and overreliance.
For a purchasing decision, translate that into a concrete data-and-failure map:
Customer records, source code, contracts, employee data, unpublished research, health information, public content, or synthetic test material.
Prompts, attachments, embeddings, logs, model-training opt-ins, feedback, outputs, and backups.
Vendor personnel, subprocessors, workspace administrators, connected apps, plugins, agents, and downstream recipients.
Wrong facts, silent omissions, harmful actions, leaked secrets, unauthorized messages, bad code, discriminatory decisions, or fabricated sources.
Ground-truth tests, citations, automated checks, human review, audit logs, monitoring, and incident channels.
Drafts are easier to recover than payments, deletions, customer communications, or regulated decisions.
Do not accept “enterprise-grade” as an answer. Ask for the specific contractual and technical controls that apply to your plan: data-use terms, retention settings, encryption, regional processing, subprocessors, audit logs, role-based access, single sign-on, export, deletion, incident notification, and the ability to disable risky integrations.
Build a real evaluation set
A model benchmark is not your business evaluation. The useful test set comes from actual work and includes the difficult cases that demos avoid.
- Sample representative tasks. Include common, long, ambiguous, multilingual, incomplete, and edge-case inputs.
- Preserve a hidden holdout. Vendors and internal prompt designers should not optimize against every test example.
- Create a scoring rubric. Separate factual correctness, completeness, policy compliance, tone, format, traceability, and action safety.
- Weight errors by severity. A harmless wording issue is not equivalent to an invented refund, exposed secret, or wrong medical instruction.
- Measure variance. Run important tasks more than once. A workflow that succeeds nine times and fails catastrophically once may be unacceptable.
- Test adversarial and messy input. Include prompt injection, conflicting instructions, unsupported files, malformed data, and irrelevant context.
Score both the raw output and the final reviewed artifact. The difference is the cost of using the tool. A system that drafts quickly but requires line-by-line verification may shift work rather than remove it.
| Metric | What it reveals | Common measurement mistake |
|---|---|---|
| First-pass acceptance | How often output can move forward without material edits | Counting any generated draft as success |
| Severe-error rate | Frequency of outcomes that could create legal, financial, security, or reputational harm | Averaging severe failures together with cosmetic defects |
| Review minutes | Human effort needed to validate and correct the output | Measuring generation speed only |
| Coverage | Share of the real workflow the tool can handle reliably | Testing only ideal inputs |
| User override rate | How often workers reject, bypass, or redo the recommendation | Treating logins or prompt counts as adoption |
| Downstream outcome | Customer resolution, conversion, defect escape, cycle time, or another business result | Stopping at model-quality scores |
Calculate the full economics
Seat price is usually the easiest number and rarely the full cost. Build a unit model around the workflow.
Licenses, API tokens, storage, premium connectors, overage, and support.
Identity, connectors, retrieval, permissions, prompt/version management, and workflow changes.
Human validation, correction, escalation, and duplicated work when users do not trust the output.
Incidents, compliance review, remediation, customer notification, and business interruption.
Compare the total against the current baseline and against a non-AI alternative. Sometimes a better form, template, search index, rule-based automation, or focused browser utility solves the same problem with lower variance and fewer governance obligations.
Also distinguish individual time savings from organizational output. A 2025 field experiment across 66 firms and 7,137 knowledge workers found that workers with an integrated generative AI tool spent two fewer hours on email each week and less time outside regular hours, but the researchers did not detect major shifts in the quantity or composition of tasks from individual-level provision alone. Time saved becomes business value only when processes, capacity, or output change.
Test integration and dependency
An AI tool can pass a quality test and still fail operationally. Evaluate how it fits the systems people already use.
- Context: Can it retrieve current, authorized information without copying entire repositories or drives into a new silo?
- Action: Which tools can it call, and can permissions be narrowed to the minimum required?
- Identity: Does every action map to a person, service account, or agent identity?
- Observability: Can administrators inspect prompts, retrieved sources, tool calls, approvals, errors, and final actions?
- Versioning: Can you identify which model, prompt, knowledge source, and connector produced a result?
- Portability: Can prompts, evaluations, logs, embeddings, and workflow definitions be exported?
- Fallback: What happens when the provider is unavailable, changes a model, removes a feature, or raises the price?
Dependency is not automatically bad. Every useful system creates some. The question is whether the dependency is proportional to the value and whether the company can continue operating during an outage or migration.
Govern permissions and accountability
A pilot that works under a founder’s account may fail when extended to hundreds of people. Governance must be part of the product test, not a later procurement exercise.
At minimum, define approved use cases, prohibited data, ownership of prompts and workflows, role-based permissions, human approval points, logging, incident escalation, retention, and periodic reevaluation. Agents that can take action require stronger controls than tools that only draft text. Sensitive actions should be explicit, reviewable, and reversible where possible.
NIST’s Govern–Map–Measure–Manage structure is useful here:
Assign owners, define risk tolerance, train users, and document approved uses.
Identify affected people, data, dependencies, intended benefits, and foreseeable misuse.
Evaluate performance, bias, privacy, security, robustness, and human-AI interaction on representative tasks.
Prioritize risks, deploy controls, monitor incidents, and stop or redesign uses that exceed tolerance.
Require named approval for financial, legal, employment, safety, or external communication decisions.
Model, connector, policy, or data changes can invalidate earlier test results.
Use a weighted scorecard
| Dimension | Weight | Pass condition | Red flag |
|---|---|---|---|
| Workflow outcome | 25% | Improves cycle time or output without degrading downstream results | Only generation speed improves |
| Quality and severity | 20% | Meets rubric and stays below severe-error threshold | Rare but consequential silent failures |
| Data and security | 15% | Controls match the sensitivity of inputs and actions | Unclear training, retention, or subprocessor terms |
| Human effort | 10% | Review and correction time falls | Workers spend more time checking fluent output |
| Integration | 10% | Fits identity, tools, permissions, and logging | Shadow accounts and broad tokens |
| Economics | 10% | Total unit cost beats the baseline with realistic volume | ROI depends on ignoring rework or low usage |
| Portability | 5% | Data and workflows can be exported or recreated | Critical process locked in opaque proprietary state |
| User adoption | 5% | Target users choose the tool and understand its limits | Mandated usage with frequent bypass or copying |
Weights should change by use case. In regulated or safety-sensitive work, severe-error risk and auditability may dominate. In low-stakes creative work, speed and user preference may matter more.
Run a 30-day pilot
Document the current process, choose users, classify data, finalize the rubric, and define stop conditions.
Run the hidden task set, compare vendors or configurations, and inspect severe failures before live use.
Deploy to a narrow group with approval gates, support, logging, and a parallel fallback process.
Collect quality, review time, override, incident, and downstream-outcome data—not only satisfaction.
Expand, revise, pause, or reject against criteria agreed before the pilot. Record what changed and when reevaluation is required.
Re-test when models, prompts, connectors, policies, or source data change.
When the correct answer is no
Reject or postpone the tool when the workflow has no measurable objective, the input data cannot be governed, severe errors cannot be detected before harm, the product requires broader permissions than the benefit justifies, the vendor cannot explain data use, or the total review burden exceeds the savings.
Also reject the false choice between “adopt this product” and “do nothing.” Alternatives include narrowing the workflow, using local processing, improving documentation, adding deterministic validation, separating low-risk drafting from high-risk approval, or building a smaller internal tool.
Not every productivity problem needs a generative model. Jivaro’s browser apps handle document comparison, text extraction, formatting, image processing, code testing, and other bounded tasks locally. They are useful controls in a broader AI workflow because they can prepare or verify material without sending it through another model.
What changes next
As model capabilities converge and change rapidly, buyers will care more about evaluations, integration, permissions, and observed business outcomes.
Identity, scoped credentials, approval gates, and tool-call logs will matter as much as prompt quality for systems that can act.
Enterprise products will compete on test management, audit, monitoring, and policy enforcement rather than a chat interface alone.
Tool sprawl, overlapping licenses, and uncontrolled data flows will push organizations toward governed platforms plus a limited set of specialized products.
Poor context, unclear ownership, weak review, and unchanged incentives will remain common barriers even as models improve.
Rapid vendor and model changes will make export, portability, and fallback architecture normal procurement requirements.
Frequently asked questions
Long enough to capture representative work and repeated use. Thirty days is a practical starting point for a bounded workflow, but seasonal, rare, or regulated processes may require a longer evaluation.
The downstream workflow outcome. Time saved matters only if quality, capacity, customer results, or cost improve after review and rework are counted.
Benchmarks can inform technical capability, but they do not replace a task set built from the company’s own workflow, data, policy, and error severity.
Classify it before testing, use synthetic or redacted examples where possible, confirm contractual and technical controls, restrict access, and avoid consumer accounts for regulated or proprietary material.
Treat that as evidence of unmet workflow demand. Provide a safe reporting path, identify the jobs people are trying to complete, then offer governed alternatives rather than relying on prohibition alone.
Build when the workflow is strategically differentiating, requires proprietary integration or controls, and can support ongoing evaluation and maintenance. Buy when the process is common and a credible vendor already meets the evidence and governance requirements.
Sources and references
- NIST: Artificial Intelligence Risk Management FrameworkNIST · reference
- NIST: Generative AI Profile (NIST AI 600-1)NIST · reference
- NBER: Generative AI at WorkNational Bureau of Economic Research · reference
- Harvard Business School: Navigating the Jagged Technological FrontierHarvard Business School · reference
- METR: Measuring AI impact on experienced developer productivityMETR · reference
- NBER: Shifting Work Patterns with Generative AINational Bureau of Economic Research · reference
- U.S. Census Bureau: The Microstructure of AI DiffusionU.S. Census Bureau · reference

