AI vendor due diligence is less about the model and more about the company and product wrapped around it. To assess an AI vendor, prove the product on your own data in a time-boxed pilot, then check, in writing, what happens to your data, how the product is secured, which models sit underneath and how changes reach you, what reliability and cost you are committing to, and how you would leave. This checklist gives IT, security, procurement and forward deployed engineering teams the questions for each area, a scoring template, red flags and a worked example.
It covers three kinds of purchase: standalone SaaS AI tools, AI features added to software you already license, and model providers you call through an API. If you are choosing between models rather than products, start with how to choose an LLM for enterprise.
Why AI vendor assessment differs from normal procurement
Your existing third-party risk process still applies. What it misses is that an AI product's behaviour is not fixed when you sign. The vendor can swap the underlying model, change a system prompt or retune retrieval, and the product that passed your pilot behaves differently next quarter. Your data may also flow to sub-processors you have never heard of, including the model provider behind the vendor.
So third-party AI risk deserves its own section in your vendor questionnaire, owned jointly by security, legal, the business owner and an engineer who can actually test the product. Results should feed your enterprise AI governance inventory, so every AI product in use has an owner, a risk tier and a review date.
Need -> Shortlist -> Questionnaire
-> Pilot on your eval set -> Score
-> Contract + exit plan -> Rollout
-> Monitor -> Re-assess on change
1. Business fit and proof on your data
A vendor demo runs on data the vendor chose, with prompts the vendor tuned. Treat it as an introduction, not evidence. The proof is a pilot against an evaluation set you built: real tickets, documents or queries, with expected outcomes agreed by your domain experts and the edge cases you know are hard. Our guide to LLM evaluation explains how to build and score that set.
- Which business process does this improve, and which metric will tell us it worked?
- Will you run a time-boxed pilot on our data, in a dedicated tenant, scored against our evaluation set?
- What integration and data preparation does the pilot need from us?
- Can we see failure cases from other deployments, not just success stories?
- How does the product handle the languages our users write in, including mixed-language input?
- What does the product do when it does not know the answer?
2. Data handling
This is where most deals should slow down. Get answers in the contract or data processing agreement, not a sales email.
- Training: Is our data (prompts, outputs, files, feedback) ever used to train or improve your models or a third party's? Is that off by default and contractual?
- Retention: What do you store (prompts, outputs, embeddings, logs), for how long, and can we configure it? Is anything kept for abuse monitoring, and who can see it?
- Sub-processors: Who are they, including model providers and hosting? How are we notified of changes, and can we object?
- Residency: Where is our data stored and processed, including inference? Can we restrict it to a region? Does support access or failover cross borders?
- Deletion: How do we delete our data, including embeddings, indexes and backups, and what confirmation do you give?
- Permissions: If the product indexes our documents, does it enforce source-system permissions at query time?
If you process personal data of people in India, the DPDP Act makes you the Data Fiduciary and the vendor a Data Processor acting for you, so your obligations follow the data into their platform. See our guide to the DPDP Act for AI applications.
3. Security
Ask for independent assurance first, then probe the AI-specific risks that certifications do not fully cover.
- Attestations: Do you have a SOC 2 Type II report (which tests controls over a period, not at a single point) or ISO/IEC 27001 certification, and does the scope include the AI product and its infrastructure?
- AI management: Are you certified to, or working towards, ISO/IEC 42001, the management system standard for AI? Our explainer on ISO/IEC 42001 covers what it does and does not prove.
- Penetration testing: When was the last independent pen test, did it cover the AI features, and can we see a summary of findings and fixes?
- Prompt injection: How do you defend against direct and indirect injection, such as instructions hidden in a document the product reads? Do you red-team the product? Our AI red teaming guide shows what good looks like.
- Agent actions: If the product can act (create tickets, reset passwords, send email), can we scope its permissions and require approval for high-risk actions?
- Identity: Do you support SSO with our identity provider and SCIM provisioning, so leavers lose access automatically?
- Audit logs: Do logs capture who asked what, what was answered, which sources were used and which actions were taken? Can we export them to our SIEM?
4. Model transparency
Many AI products are a layer on someone else's model. That is fine, but the model provider becomes part of your supply chain, so you need to know.
- Which models does the product use, from which provider, hosted where? Is there a fallback model during outages?
- Will you notify us before changing a model, model version or major prompt or retrieval logic? How much notice?
- Can we pin a version, or opt into changes after re-running our evaluation set?
- For model providers: what is the deprecation policy, and how long is a version supported after a successor ships?
5. Quality evidence and evaluation access
- How do you evaluate quality internally, on what data, and how often? Can you share the method?
- Do you run regression tests before each release?
- Can we run our own evaluation set against the product through an API, repeatedly, after go-live?
- Does the product show sources so users and reviewers can verify answers?
- How are wrong answers reported, triaged and fixed?
Evaluation access after go-live is the question vendors most often dodge, and the one that matters most. Without it you cannot detect when an upstream model change degrades your use case.
6. Reliability, SLAs and rate limits
- What uptime commitment applies to the AI features specifically, and what are the service credits?
- What rate limits or usage quotas apply per user and per tenant, and what happens when we hit them: queueing, errors or a silent downgrade to a smaller model?
- What latency should we expect at our volume, from our region?
- How do you handle an outage at your model provider, and is the fallback tested?
- Is there a degraded mode where the product keeps working without AI?
7. Cost model and lock-in
AI pricing comes per seat, per conversation or resolution, per token or credit, or bundled into a higher licence tier. Model your expected usage under each, including growth, because consumption pricing can surprise you once adoption takes off.
- What exactly is the billing unit? Are retries, failed calls and admin testing billed?
- Can we set usage caps and alerts?
- Can we export configuration, prompts, knowledge sources, conversation history and feedback in a usable format?
- Are embeddings or fine-tuned models built from our data portable?
- Do integrations use open standards, such as the Model Context Protocol for tools, so our work is reusable?
Lock-in becomes a problem when you discover it at renewal. Ask the export questions while you still have leverage.
8. Legal and regulatory
Bring legal in early. These are topics to raise, not legal advice.
- IP indemnity: Does the vendor indemnify you if output infringes a third party's intellectual property? Many providers offer some form of this, but conditions vary, often requiring built-in filters to stay on. Read the conditions.
- Liability: What is the cap, and are data and confidentiality breaches treated separately?
- Ownership and confidentiality: Who owns outputs, and does confidentiality explicitly cover prompts, outputs and files?
- Regulatory support: Will the vendor help with DPDP requests, sector regulator audits and breach notification within your timelines?
- EU AI Act: If the system serves people in the EU or its outputs are used there, ask how the vendor classifies it, what documentation it provides as a provider, and what it expects of you as a deployer. See the EU AI Act for Indian IT teams.
9. Responsible AI
- Have you tested for bias across groups relevant to our use case (language, region, gender, role)? Can you share the method? Our guide to bias and fairness testing for AI covers what a credible answer includes.
- Which oversight features exist: review queues, approval steps, confidence thresholds, easy escalation to a person? See human-in-the-loop AI for the patterns.
- Are users clearly told when they are reading AI-generated content?
- Do you publish documentation of intended use and known limitations?
10. Exit plan
Write the exit plan before you sign: what you export, in which format, how deletion is certified, how long the transition lasts and what covers the process meanwhile. Keep the business process documented independently of the tool.
- What transition assistance is available at termination, and at what cost?
- What happens to our data if you are acquired, change model provider or retire the product?
Building this evaluation discipline into an IT or procurement team takes practice. Cloudsoft's corporate AI training for teams runs hands-on workshops where your engineers and risk owners build an evaluation set and pilot plan for a tool on your shortlist.
AI vendor assessment scoring template
Separate gates from scored criteria: a vendor that fails a gate is out, however good the demo. Set weights before you see any results.
| Area | Type | Evidence | Weight | Vendor A | Vendor B |
|---|---|---|---|---|---|
| Quality on our evaluation set | Scored, minimum bar | Pilot results scored by our experts | High | ||
| No training on our data | Gate | Contract or DPA clause | Pass/fail | ||
| Residency and sub-processors | Gate | Sub-processor list, region settings | Pass/fail | ||
| Security assurance | Gate | In-scope attestation, pen test summary | Pass/fail | ||
| SSO, SCIM, audit logs | Gate | Tested in pilot tenant | Pass/fail | ||
| Injection and action controls | Scored | Our red-team results in pilot | High | ||
| Model change notice and pinning | Scored | Contract notice period | Medium | ||
| Evaluation access after go-live | Scored | API access demonstrated | Medium | ||
| Reliability and limits | Scored | SLA terms, load test | Medium | ||
| Total cost at projected usage | Scored | Our cost model | Medium | ||
| Export and portability | Scored | Sample export inspected | Medium | ||
| Legal and responsible AI | Scored | Legal review, oversight features tested | Medium |
Use a simple scale with written definitions for each level, and record the evidence behind every score.
Red flags
- The vendor refuses a pilot on your data, or runs it with no access for your team.
- "We don't train on your data" appears in marketing but not the contract.
- No sub-processor list, or the model provider is not named.
- The security attestation covers a different product or only the corporate environment.
- No answer on prompt injection beyond "the model is safe".
- Model changes happen silently, with no notice or version pinning.
- Audit logs cannot be exported.
- Accuracy claims with no method behind them.
- Data export is unavailable or sold only as professional services.
One red flag is a negotiation point. Several together usually mean the product is not ready for enterprise use.
Running a fair bake-off
- Freeze the evaluation set and weights first, and keep a held-out portion no vendor ever sees.
- Give every vendor the same data, time and access, including the same number of configuration days.
- Score blind where possible, so reviewers do not know which product produced an answer.
- Test the failure paths: injected instructions, out-of-scope questions, a user without permission to the source, an upstream outage.
- Price the same usage scenario for every vendor.
- Write up the decision with the evidence, the trade-offs accepted and the triggers for re-assessment.
This is the work forward deployed engineers do from the other side of the table: proving an AI product inside a customer's environment, with the customer's data and controls. If that appeals as a career, Cloudsoft's Forward Deployed Engineer course teaches it through a simulated customer engagement.
Illustrative example: a GCC buying an AI service-desk tool
Consider a global capability centre in Hyderabad that runs the IT service desk for its parent company's employees across several countries. It wants an AI assistant to answer common questions, triage tickets and handle simple requests such as access and software installs.
Shortlist and gates. Two products make the shortlist: the AI add-on for the existing ticketing platform and a standalone SaaS assistant. The parent's data protection office requires approved regions and no training on employee data. The standalone vendor's contract allows "service improvement" use of conversations by default; it agrees to an opt-out clause. The add-on passes on residency but cannot yet export audit logs to the SIEM, so the team records the vendor's committed roadmap date as a condition.
Pilot. The evaluation set comes from anonymised past tickets: access issues, VPN problems, laptop requests and country-specific policy questions. Adversarial cases include a knowledge article with a hidden instruction to grant admin rights, a user asking about a confidential HR policy they cannot open, and questions mixing English and Hindi. Both products run in sandbox tenants over SSO for the same period.
Findings. The standalone assistant answers more questions correctly, but it summarised a restricted document for an unauthorised user, because its connector indexed content with a service account. The add-on scores lower on answer quality but inherits the platform's permissions, and its password-reset action requires approval. The team treats the permission leak as a gate failure until it is fixed and re-tested.
Decision. The GCC starts a limited rollout of the add-on, with conditions: audit log export by the committed date, notice before model changes, re-running the evaluation set after any model change, and an exit plan that keeps the knowledge base in the ticketing platform rather than inside the AI feature. For how such an assistant is engineered, see our Microsoft Teams IT helpdesk AI assistant project.
Frequently asked questions
What is AI vendor due diligence?
AI vendor due diligence is assessing an AI product or provider before you buy it. It covers business fit proven on your own data, data handling, security, model transparency, quality evidence, reliability, cost and lock-in, legal terms, responsible AI and an exit plan.
How is it different from a normal vendor security questionnaire?
A normal questionnaire assumes the software behaves the same after you buy it. AI products change when the vendor swaps a model or prompt, outputs vary, and data may flow to model providers as sub-processors. AI due diligence adds questions on training use, model changes, prompt injection, evaluation access and output quality.
Which certifications should we ask an AI vendor for?
Ask for an independent security attestation such as a SOC 2 Type II report or ISO/IEC 27001 certification whose scope covers the AI product. ISO/IEC 42001 certification shows the vendor runs an AI management system. Certifications are a baseline, so still test AI-specific risks in the pilot.
How do we stop a vendor training on our data?
Put it in the contract or data processing agreement, not just a settings page. The clause should cover prompts, outputs, files and feedback, bind the vendor's model providers too, and apply by default to your tenant.
Do AI features inside software we already use need due diligence?
Yes. An AI add-on can introduce new sub-processors, retention, permissions and behaviour inside a product you already trust. Run the same checklist, focusing on what the AI feature changes compared with your existing contract.
How often should we re-assess an AI vendor?
Annually, and on triggers: a model change, a new sub-processor, a pricing change, a security incident, a new use case or a regulatory change. Re-run your evaluation set whenever the underlying model changes.
Who should own AI vendor assessment?
No single team. Procurement runs the process, security and privacy own the data and security gates, legal owns the contract, the business owner defines success, and an engineer runs the pilot. Record the decision in your AI governance inventory.
If your IT, security or procurement teams are about to buy AI tools and want a shared, practical way to evaluate them, Cloudsoft's corporate training programs can be tailored to your shortlist and controls, delivered at your site, in our Ameerpet classroom or live online. Call +91 96660 19191 to arrange a free demo session.



