What should you ask before hiring an AI consulting agency?

Use these questions to ask AI consulting agency candidates for evidence about the proposed workflow, not a general tour of their capabilities. Start with the business decision, then test delivery proof, data handling, controls, operating cost, and exit terms. Zyphh's workflow-first approach is one example of the position this checklist is designed to examine, not an exemption from the same scrutiny.

Send the questions before the call and ask every candidate about the same representative workflow. A useful answer names an artifact you can inspect, a person who owns it, and a condition that could change the plan. A confident adjective, polished demo, or list of model names is not evidence on its own.

Which questions test problem fit and delivery proof?

The first four questions show whether the agency understands the work and can support its claims. GOV.UK's AI procurement guidance recommends a clear problem statement, openness to non-AI alternatives, data assessment, iterative delivery, and lifecycle planning. That is a sound starting point outside public procurement too.

  1. 1. What workflow and business outcome are we changing? Ask for the trigger, current path, baseline, owner, target measure, and definition of done. A proposal that starts with a chatbot, agent, or model before it defines the work has reversed the decision.
  2. 2. Why does this need AI? Ask which parts require probabilistic interpretation and which should use ordinary rules, product configuration, search, or a person. A credible agency can recommend a smaller non-AI fix.
  3. 3. What relevant delivery evidence can we verify? Request a client reference when permission exists. If confidentiality prevents that, ask for redacted scopes, test plans, architecture decisions, operating runbooks, or post-launch review examples. Check that the evidence matches your type of workflow, not just your industry label.
  4. 4. Who will do the work? Name the people responsible for discovery, data, integration, security, testing, and post-launch operation. Confirm how much of their time is committed and whether the sales team hands the project to someone else.

What should you ask about data, controls, and testing?

These questions turn responsible-AI language into operating requirements. The NIST Generative AI Profile points buyers toward acquisition due diligence, service agreements, third-party transparency, and documented test, evaluation, validation, and verification before deployment.

  1. 5. Where will our data go? List every model provider, cloud service, subcontractor, log store, and support tool that may receive it. Ask about access, encryption, region, retention, deletion, backups, model training, tenant separation, and incident notice. The ICO's AI contract guidance calls for whole-supply-chain due diligence, written responsibilities, subprocessor controls, deletion or return terms, and continuing review.
  2. 6. Which actions can the system take? Require a list of read, draft, send, approve, write, and delete permissions. For each consequential action, identify the approval rule, log record, rollback path, and person accountable for the result.
  3. 7. How will you prove the system works for our workflow? Ask for representative test cases, expected outputs, acceptance thresholds, known limits, and results split by important customer or workflow groups. A generic model benchmark does not prove reliability in your data, integrations, or operating context.
  4. 8. What happens when a dependency fails? Walk through stale data, missing fields, duplicate events, rate limits, unavailable APIs, model changes, and low-confidence output. The answer should cover detection, retries, safe stopping, alerts, manual fallback, recovery, and incident review.

Which questions expose cost and lock-in?

The last two questions test whether the buyer can operate, change, and leave the system. CISA's Secure by Demand guide gives software customers questions for product security before, during, and after procurement. Apply the same continuing review to the custom code and third-party components an agency selects.

  1. 9. What is the complete operating cost? Separate discovery, build, integration, data cleanup, internal time, hosting, model and API use, monitoring, security review, support, maintenance, retraining or prompt changes, and exit work. Ask which assumptions make the estimate rise.
  2. 10. What do we receive if the engagement ends? Put ownership and access in the contract. The handoff may need code, configurations, prompts, evaluation sets, infrastructure definitions, credentials, documentation, logs, data exports, licenses, open issues, training, and a transition period. Clarify what remains proprietary and whether another team can operate the result.

How should you compare AI consulting agencies?

Score the evidence immediately after each call, before price or rapport blurs the gaps. Use 0 for no evidence, 1 for a partial or promised answer, and 2 for a relevant artifact or contract term your team can inspect. Do not average away a zero in data protection or a control your risk owner considers mandatory.

AreaEvidence worth requestingCommon red flag
Problem fitWorkflow boundary, baseline, outcome, alternativesSolution chosen before discovery
Delivery proofRelevant reference or redacted operating artifactDemo presented as business proof
Data and controlData-flow map, permission list, logs, rollbackVague claim that the system is secure or compliant
Testing and operationTest set, thresholds, failure drill, monitoring planOne accuracy number with no context
Cost and exitLifecycle estimate, handoff inventory, termination termsOwnership that still requires permanent vendor access

Ask finalists to apply the checklist to one sanitized, difficult case rather than a friendly demo. Record what each candidate would refuse to automate and what evidence could make it change that decision. Good judgment is easier to see at the boundary than in the happy path.

Which answers should pause the buying process?

Pause when an agency guarantees an outcome it cannot control, quotes a production build without examining the workflow or data, treats all human review as temporary friction, or cannot explain where data and automated actions go. Also pause when the contract leaves testing, maintenance, model changes, security response, or handoff to later discussion.

A small firm should not lose automatically because it lacks a long client list, and a large firm should not pass because its logo deck is familiar. Adjust the depth of diligence to the consequence of the use case. For any candidate, use the same evidence standard and document exceptions.

Before a pilot starts, turn the chosen answers into acceptance criteria. The AI automation guardrails guide can help translate permissions, approvals, logs, stop rules, and recovery into a reviewable operating plan.

Turn this into your own build plan.

Run the Workflow Opportunity Score or book a strategy call. Bring one repeated workflow, the tools involved, and the number that should move.

Run the score

Sources and further reading

  1. NIST: Artificial Intelligence Risk Management Framework, Generative AI Profile
  2. CISA and FBI: Secure by Demand Guide
  3. GOV.UK: Guidelines for AI procurement
  4. UK ICO: AI contracts and third parties

FAQ

What is the most important question to ask an AI consultant?

Ask what workflow will change, how success will be measured, and what would make the team stop. That answer reveals whether the consultant is diagnosing a business problem or fitting your company to a preferred tool.

Should an AI consulting agency share client work?

Ask for permissioned references or evidence relevant to your workflow. Confidentiality may rule out names and raw data, so accept redacted scopes, test plans, runbooks, or architecture decisions when they let your team verify how the agency works.

Who should help evaluate an AI consulting agency?

Include the workflow owner plus the people responsible for data, security, and operating the finished system. Add legal, privacy, procurement, finance, or domain specialists when the data, decision, contract, or consequence requires them.

Is a paid discovery phase a red flag?

No. A bounded discovery phase can reduce risk when its deliverables, price, decision point, and ownership are explicit. It becomes a concern when discovery has no clear output or automatically commits you to the build.