In short: Choosing a first AI use case works best with a feasibility-first approach: start with a project you can actually ship, prove the value, and let momentum carry you to the bigger ones.
Most enterprises are choosing their first AI project backwards. They start from what the technology can do at its most impressive, then work toward whether they can run it, when the order that actually ships things runs the other way. The results show up in the data. MIT’s Project NANDA puts the share of enterprise generative AI pilots that deliver no measurable return at about 95%. S&P Global found 42% of companies abandoned most of their AI initiatives in 2025, more than double the prior year. These failures have less to do with the models than with selection: tools that never reach a real workflow, projects with no baseline to measure against, owners who run innovation labs rather than the process the tool was meant to fix.
Where Pilots Go Wrong
A pattern connects most stalled pilots. The project was chosen for its ceiling rather than its floor. Someone asked what the most transformative application of AI could be, when the more useful question was which application the organization could actually operate: with its current data, its current processes, its current governance maturity, and its current tolerance for error.
A Framework for Picking the First Project
The most durable selection frameworks reduce to two honest axes. One measures business value. The other measures feasibility, and feasibility is where most selection goes wrong. A feasible project needs more than a capable model. The data has to exist, the workflow has to absorb the tool, the oversight has to be staffed, and the governance requirements have to be met. That broader definition is baked into the major frameworks: PwC and Google both score value against feasibility, and weighted models add data readiness, risk, and compliance as criteria in their own right.
Scoring Feasibility Honestly
Feasibility assessment breaks down the moment it turns into a sales pitch. Teams that want a project tend to score it generously. Four criteria keep the exercise grounded.
Data readiness. The project should run on data the organization already has, already governs, and already trusts. A use case that depends on a six-month data cleanup effort is a data project wearing an AI costume, and it should be scoped and staffed as one.
Workflow integration. The MIT research was blunt on this point: tools that require people to change how they work get abandoned, while tools that slot into existing workflows get used. If adopting the system requires a new interface, a new habit, and a new approval chain, the odds lengthen considerably.
Human oversight. A first deployment should keep a person between the AI output and any consequential action. Beyond catching errors, the review step shows the organization how the system actually fails, which is knowledge no vendor deck provides.
Contained blast radius. When the system produces a wrong answer, and it will, the damage should be internal, recoverable, and educational. An error in an internal knowledge search result costs a few minutes. An error in a public commitment costs a tribunal hearing.
Value scoring has its own failure mode: projecting transformation instead of measuring improvement. Grounded value has three markers. The task is repeatable, happening dozens or hundreds of times a week, so small per-instance gains compound. The current process has a measurable cost in hours, dollars, or delay, so improvement can be demonstrated rather than asserted. And the process has an owner who feels the pain, so adoption has an internal champion with a stake in the outcome.
Where to Start: Four Proven Patterns
The archetypes that consistently clear both axes share a profile: internal-facing, grounded in the organization’s own content, and structured with a human in the loop.
Internal knowledge search and retrieval. Retrieval-augmented systems that answer employee questions from the organization’s own documentation are the closest thing to a canonical first project. The data already exists. The blast radius is internal. Answers can cite their sources, which makes verification natural. And the baseline, time spent hunting for information, is easy to measure before and after.
Content tagging and metadata automation. Large content operations accumulate years of inconsistently tagged material. AI-assisted classification, with editors approving suggestions rather than typing taxonomies from scratch, delivers compounding value: better search, better recommendations, better analytics. For organizations running enterprise WordPress at scale, where content libraries routinely run to tens of thousands of items, this is often the highest-leverage unglamorous project available.
Semantic site search. Replacing keyword matching with meaning-based search improves an existing feature rather than introducing a new behaviour. Users already search. Nothing about their workflow changes. The improvement shows up directly in search success metrics, which most analytics stacks already capture. Vector-based retrieval over a structured content repository is a mature pattern at this point rather than a research problem.
Editorial workflow assistance. Draft summaries, headline variants, accessibility descriptions, and internal link suggestions, all surfaced inside the tools editors already use, with every output passing through human review before publication. The publishing workflow itself becomes the oversight mechanism, which is why content operations make such a natural proving ground for enterprise AI generally.
None of these projects would headline a conference keynote, but all of them ship, get used, and generate the measured evidence that makes the second project easier to approve.

From First Project to Second
The first deployment sets the pattern for everything after it, so the sequence matters as much as the selection. It starts with the baseline: measuring the current process before the AI touches it, because improvement claims without a baseline are testimony rather than evidence. Then a scoped deployment with defined users, defined success metrics, and a defined review period. Then honest measurement, including the failures, since the failure modes discovered on a low-stakes project are the cheapest education an organization will ever get in AI operations. Then, and only then, expansion, with the governance practices, oversight habits, and measurement discipline from the first project carried forward as institutional capability rather than reinvented from scratch.
Organizations that follow this arc end up with something more valuable than a single working tool. They end up with a repeatable method for turning AI candidates into AI deployments, and that method outlasts any individual tool it produces.
The Canadian Layer
Canadian enterprises choose first projects inside a regulatory environment that has been moving quickly, and governance readiness belongs on the feasibility axis rather than in a compliance appendix.
Quebec’s Law 25 requires transparency around automated decision-making that affects individuals, along with meaningful explanation on request. Ontario’s Employment Standards Act amendments, in force since January 2026, require employers to disclose the use of AI in screening job applicants. OSFI’s Guideline E-23 brings model risk management expectations to federally regulated financial institutions, with AI models squarely in scope. At the federal level, AIDA died on the order paper, but Ottawa has continued building out its AI strategy, including safety-institute work, while sector regulators fill the gap in the meantime.
For use case selection, a first project that avoids personal information, avoids automated decisions about individuals, and stays internal faces a dramatically lighter governance burden than one that touches any of those lines. This is a sequencing argument rather than a case for avoiding regulated use cases altogether. Build governance muscle on a project where the stakes are low, then carry that muscle into the projects where the stakes are real. Organizations that invert the order end up learning privacy law and AI operations simultaneously, under pressure, in public.
Statistics Canada’s surveys have shown Canadian business AI adoption rising but still uneven, concentrated in larger firms and specific sectors. The organizations that get a first deployment into production, however modest, are already ahead of a surprising share of their market.
The Deployable Project Wins
The gap between the 95% and the 5% comes down to selection discipline rather than talent, budget, or model access. The organizations getting real returns from AI picked first projects that were boring on purpose: internal, measurable, grounded in existing data, wrapped in human review, and clean from a governance standpoint. Then they let the evidence do the persuading. The impressive projects can come later, once there is a working deployment underneath them.
Where a Consulting Partner Fits
Most of the discipline described above is organizational rather than technical, which is why the hardest part of a first AI project is rarely the build. It is the honest scoring, the baseline measurement, and the governance groundwork that internal teams, close to their own projects, find difficult to do objectively.
Trew Knowledge’s AI consulting practice covers that full arc. Strategy engagements assess readiness and produce a roadmap tied to business goals. LLM and agent development builds the project around workflows teams already run, which is where adoption is won or lost. Data and infrastructure enablement closes the readiness gaps assessment uncovers, managed AI operations keeps performance and compliance measurable after launch, and our generative AI practice adds production capability under the same human-review discipline.
It is no accident that the archetypes earlier in this article live in content operations. We come from that world, with over a decade building and running enterprise publishing platforms, so we know where the deployable projects hide inside the workflows organizations already run. Talk to our team about finding the first project worth deploying, and where it likely already exists in your organization.
