How to Hire an AI Consultant: A Step-by-Step Process for Business Leaders

Step one is not a vendor search. It is a sentence. Before you contact anyone, write down the business outcome you are buying, in the form: we want to move this metric, for this process, owned by this person, by this date. “Reduce average claim intake handling time from nine minutes to four minutes for the commercial lines team by the end of Q3” is a brief. “Explore AI opportunities” is a budget with no owner.

That sentence does most of the work in the hiring process that follows. It tells you what kind of engagement you need, it tells vendors what to price, and it gives you the acceptance criteria you will eventually put in the contract. If you cannot write it yet, that is useful information too. It means what you should buy first is a short paid discovery, not a build.

Decide what kind of engagement you are actually buying

Most bad AI vendor relationships start as a mismatch between what the buyer needed and what the vendor sells. There are four common shapes, and they are not interchangeable.

Strategy and roadmap

A short engagement, usually a few weeks, that produces a prioritized set of use cases, an honest data readiness assessment, and a build versus buy recommendation for each item. Buy this when you have six ideas floating around the leadership team and no way to rank them. Do not buy this when you already know exactly what you want built. Paying for a roadmap you could have written yourself is the most common form of wasted first spend.

Build

A defined deliverable shipped into production. This requires a specific use case, accessible data, and a named internal owner. Vendors who are good at this will push back on your scope before they agree to it.

Embedded team

Their engineers, your roadmap, your project management. This works when you have direction and no capacity, and when you have someone internal capable of directing technical work day to day. If you do not, an embedded team drifts.

Managed service

Someone runs the system after launch: monitoring, evaluation, retraining, incident response, cost management. AI systems degrade quietly. Data shifts, upstream systems change, model providers deprecate versions. Decide who owns that before you go live, not after.

In practice, many buyers need a small paid discovery followed by a build. Contract those separately, with a decision gate between them. A vendor who insists on a single large signature covering both is optimizing for their pipeline, not your risk.

Write a scope vendors can price consistently

If three proposals come back and you cannot compare them, the problem is usually your scope document. Give every vendor the same package, and include:

  • The business outcome and the metric, with the current baseline if you have it
  • The current process, described step by step, including the manual workarounds
  • Systems of record involved, and how data would be accessed (API, warehouse, export)
  • Data volume, history depth, and sensitivity classification
  • Who owns the decision internally, and who the subject matter experts are
  • Success criteria for a pilot, stated as a threshold, not a feeling
  • Production requirements: users, latency, uptime, languages, compliance regime
  • Timeline, and your internal constraints (security review, procurement cycles, change freezes)
  • What you will provide: SME hours per week, environments, test data, review turnaround

Then ask each vendor to price the same three things: discovery, a first production increment, and one year of run support. Comparable line items beat comparable totals.

One warning from experience: internal security review is the single most common cause of schedule slippage on these projects. Start it before kickoff, not at integration time.

Where to source candidates

Start with firms already inside your environment. Your existing systems integrator or software partner knows where your data lives and has already passed your security review, which is worth more than a polished capabilities deck. Then look at partner directories for the platforms you already run (Azure, AWS, Salesforce, Dynamics, NetSuite), peers in your industry with a similar compliance profile, and curated vendor lists such as this overview of top US-based AI consultants and developers, which is a reasonable way to build an initial long list before you shorten it against your own criteria.

Two sourcing habits to break. Do not build a shortlist from cold outbound decks. Do not shortlist on rate card alone, because a lower hourly rate applied to a team that needs three months to understand your data is not cheaper.

Screening criteria before you spend meeting time

Cut a long list to three finalists using criteria that are hard to fake:

  • Production systems shipped under data conditions like yours, not demo environments
  • Domain adjacency, or a credible explanation of how they get up to speed
  • A described practice for evaluating output quality, offered without prompting
  • Security posture and willingness to work inside your environment and accounts
  • Transparency about who staffs the work, including subcontractors and delivery locations
  • References you can actually call, including one project that changed direction
  • Demonstrated willingness to tell a prospect no

Three finalists is enough. Beyond that, your own evaluation quality drops.

The technical evaluation

Run a working session, and pay for it. Give each finalist a sanitized data sample and a real problem, and watch how they work rather than how they present. The useful signal is not whether they can name models. It is whether they can describe how they will know if the output is good enough.

Ask them to walk through evaluation concretely: how a test set gets built, who produces the labeled ground truth, what the human review rubric looks like, how failures are categorized, what threshold justifies shipping, and how quality gets monitored after launch. A vendor who has done this work will have opinions and examples. A vendor who has not will retreat to model names and benchmark scores.

Bring your own engineer and your data owner into the room. The person who knows why three fields in the source system are unreliable will learn more in an hour than procurement will learn in a month.

An interview question set, with the answers that matter

Use these in the finalist sessions. For each, the competent answer and the overselling answer look very different.

  1. Tell me about a project where the model did not perform well enough. Competence: a specific case, a diagnosis, and what changed, including projects they recommended stopping. Overselling: no failures, only success stories.
  2. How would you measure quality for our use case? Competence: a labeled evaluation set, a named party who does the labeling, an acceptance threshold, and a distinction between offline evaluation and production monitoring. Overselling: an accuracy figure with no test set behind it.
  3. What data do you need from us, and in what form? Competence: questions about systems of record, refresh frequency, historical depth, PII handling, and known edge cases. Overselling: give us access to everything and we will figure it out.
  4. Who exactly will do the work? Competence: named people, roles, allocation percentages, locations, and any subcontractors. Overselling: senior names in the pitch who are not in the staffing plan.
  5. What happens if we want to change model providers in a year? Competence: an abstraction layer, portable prompts and evaluation assets, and a realistic estimate of switching cost. Overselling: a proprietary platform that quietly owns your workflow.
  6. What will this cost to run per month at our volume? Competence: an estimate with stated assumptions covering inference, infrastructure, human review, and monitoring. Overselling: no answer, because run cost is not their problem.
  7. What part of this should we not build? Competence: a recommendation to buy something off the shelf, or to fix a process before automating it. Overselling: everything is custom.
  8. How do you hand this over? Competence: repositories in your accounts from day one, documentation as an acceptance criterion, pairing and training built into delivery. Overselling: knowledge transfer as a line item at the end.

Reference checks worth doing

Ask for three references: one similar in scope, one where the project changed significantly, and one that has ended. The third is the one that teaches you something.

Questions that produce real answers:

  • What was the original scope, and what actually shipped?
  • Who from the vendor did the work, and did the team change mid engagement?
  • How did they handle the first missed estimate?
  • What did your internal team end up doing that was not in the plan?
  • Is the system still running, and who maintains it now?
  • What would you scope differently if you started again?
  • Would you hire them again for this same work?

Listen for hesitation on the last one. If every reference is an active engagement, ask why.

Comparing proposals and pricing models

Fixed fee works for bounded, well specified deliverables such as discovery or an integration with known endpoints. It becomes friction when requirements move, which they do on AI work, so expect change orders.

Time and materials is honest for exploratory work, but only with a not to exceed cap, weekly burn reporting, and a named person on your side watching it.

Retainers fit run and support. Define what the retainer buys in units: monitoring coverage, evaluation cycles, response times, a monthly improvement allocation.

Outcome based pricing sounds ideal and is rarely clean. It requires a baseline both sides accept, instrumentation neither side controls unilaterally, and an attribution model that survives a good quarter. Use it for narrow, well instrumented metrics. Do not use it as a substitute for a scope.

A defensible default: fixed fee for discovery, time and materials with a cap for the first build increment, retainer for run. Compare vendors on first year total cost, including the run cost they quoted, not on proposal price.

Contract terms that actually matter

  • IP ownership. Name the artifacts: source code, prompts, fine tuned weights, evaluation datasets, and any labeled data your team produced. Prompts and evaluation sets are real assets and are frequently left unassigned.
  • Model and data rights. No training on your data without written consent. Get the sub processor list, where inference happens, and retention terms from the model providers in the stack.
  • Data handling and confidentiality. Environment and access controls, PII treatment, a data processing agreement, breach notification timelines, audit rights.
  • Subcontracting and offshore disclosure. Named subcontractors, countries of delivery, and your approval right over changes to key personnel.
  • Exit and knowledge transfer. Repositories and infrastructure in your accounts from the start, documentation as an acceptance criterion rather than a promise, and a transition assistance clause at a defined rate for a defined period.
  • Acceptance criteria tied to the evaluation thresholds you agreed during the technical evaluation, not to a demo.

The first 30 to 90 days

Days 1 to 30: access granted, data profiled, current baseline measured, evaluation set built and signed off by your subject matter experts, scope confirmed or revised in writing. The deliverable is a baseline and an evaluation harness, not a prototype.

Days 31 to 60: a working system running against real data in a non production environment, weekly demos, a documented failure taxonomy, and a decision gate where continuing, changing direction, or stopping are all live options.

Days 61 to 90: the production path. Integration, monitoring, the human in the loop workflow, a runbook, first real users, and measurement against the baseline you captured in month one.

Throughout: one internal owner with decision authority, weekly written status, and no invisible work.

Red flags checklist

  • Proposes a solution before asking anything substantive about your data
  • Cannot describe how they will evaluate model output quality
  • Case studies with no measurable outcome, only technology names
  • Unwilling to run a paid discovery, or insists on one large signature up front
  • Vague about who staffs the work, or about subcontractors and delivery locations
  • No answer on ongoing run cost
  • Wants to host your data in their environment without a clear reason
  • Resists putting code and infrastructure in your accounts
  • Claims certainty about accuracy before seeing your data
  • Every internal question routes through a salesperson
  • No opinion about what you should not build

Making the call

Rank finalists on evidence, not enthusiasm: production work under similar conditions, a credible evaluation practice, staffing you can verify, and contract terms they will sign without negotiation theater. If two are close, choose the one that asked better questions about your data.

If you are still building a long list, this roundup of US-based AI consulting and development firms is a practical starting point. Run every name on it through the screening criteria, interview questions, and red flags above before you shortlist.

Questions?

We’re happy to discuss your technology challenges and ideas.