AI vendor risk assessment for shopping agents

Run a vendor risk assessment on an AI shopping agent before you sign: security reports, a data processing agreement, contract terms and a paid pilot.

A vendor risk assessment checks whether an AI shopping agent is safe and practical for your store before you sign. Start with the data it will see and the actions it may take. Then request security evidence, settle the data and contract terms, score the open risks and run a short pilot against questions your customers actually ask.

You can keep this to one decision file. Give an owner to each open risk and agree which issues must be fixed before launch. The example scorecard below is a reusable starting point, not a claim that any particular vendor passes it.

Where this fits in third-party risk management

Third-party risk management is the way a retailer checks, contracts with and keeps watch on an outside supplier whose service could affect its customers or operations. It matters here because a shopping agent may read customer messages, show product or delivery information and, if connected and authorised, act on a basket or order. A mistake can reach a customer before your team sees it.

Treat this check as one entry in your existing supplier process. NIST’s supply-chain guidance describes a cycle of identifying and assessing risk, choosing a response and monitoring changes. For a store, that becomes a short workflow:

  1. Plan: Name the store problem, the data the agent needs and the actions it may take. Limit the initial access to what the pilot needs.
  2. Check: Ask for security evidence, map data flows and test answers and actions. Record any gap with an owner and due date.
  3. Contract: Put the agreed controls, service commitments and exit steps in writing before production access.
  4. Monitor: Revisit incidents, model or sub-processor changes, test results and contract commitments at an agreed interval.
  5. Offboard: Remove access, export what you need and confirm deletion under the agreed terms.

The vendor risk assessment sits mainly in planning and due diligence, but it gives you the monitoring and exit questions as well. A small team can use a one-page vendor policy: who approves a supplier, which data and actions trigger a deeper check, the documents required, who accepts any remaining risk and when to review it. Keep the same form for every candidate. This prevents a favourite demo from changing the pass criteria halfway through procurement. Follow any organisation-wide AI policy your retailer already has; that wider policy is a separate decision.

The National AI Centre’s adoption guidance calls for assessing each AI use case, including third-party systems, against the likely harm and its impact. For a retailer, a read-only product adviser needs a different review from an agent that can change a cart or place an order. Ask who supplies the underlying model, whether chat data can reach that provider and how the vendor controls wrong answers or unauthorised actions. Those are the AI-specific parts of this third-party check.

Write the AI business case

Put the decision on one page before comparing vendors. State the customer problem in store terms, such as repeated size questions that staff answer by hand or shoppers who cannot find a delivery condition. Then record the expected gain, the full cost and the risks that the vendor must clear. The Digital Transformation Agency’s public-sector procurement guidance recommends clear outcomes, a business case and a team with technical and business input. A private retailer can use those planning steps without treating government procurement rules as its own.

A useful business case has the current baseline and a pilot target beside it. For example, if your team spends time answering delivery questions, measure the hours now, then set a target for correctly resolved questions in the pilot. For sales, use a store measure you can compare fairly, such as assisted sessions that reach checkout, with the same definition before and during the test. Include licence fees, setup, staff review time, integration work, support and expected renewal cost. A guide to chatbot return on investment can help structure that cost and savings calculation.

Write the pilot’s success measures here, before the vendor sees the test questions. Also record a stop rule, such as any wrong order action or repeated invented stock answer. The business case then tells finance what value you expect and tells IT which risks cannot be traded away for a high demo score.

Security evidence: SOC 2, ISO 27001 and penetration tests

Ask which system each document covers and when it was tested. A security badge alone cannot tell you whether the shopping agent, its model provider and its connected store data were in scope.

  • SOC 2 Type II report: An independent auditor examines the described service controls and their operation over a period. Read the system boundary, period, exceptions and any customer controls you must operate yourself. A Type I report addresses control design at a point in time; it does not give the same period of operating evidence. AICPA’s discussion of SOC reports explains why the report’s subject and type matter. Neither report proves that every future AI answer will be correct.
  • ISO/IEC 27001 certificate: This concerns the vendor’s information security management system. Check the certified organisation, sites, service scope, issuer and expiry. Ask for the statement of applicability, which records the controls selected for that system and why controls were included or excluded. ISO explains the standard and its technical committee explains the statement. A certificate covering an unrelated office or product does not settle your agent review.
  • Recent independent penetration test: This is an attempt to find exploitable weaknesses in the tested system. Ask for the scope and test dates, then read the findings by severity, planned fixes and retest results. NIST’s security-testing guide recommends reporting findings and mitigation actions. The summary can omit exploit detail and personal information while still showing whether serious findings remain open. A vendor may share a fuller report under a non-disclosure agreement. A clean test of one version cannot guarantee the next release is safe.

If the vendor has none of these yet, ask for evidence of access controls, staff permissions, encryption in transit and at rest, security monitoring, vulnerability fixes and incident notification times. Ask whether its model provider and hosting suppliers have been assessed. Record who will verify each answer and what evidence would close the gap. For an agent that can change a basket or order, request a demonstration of action permissions and launch guardrails as well as general security controls.

Data processing agreement and chat data ownership

A data processing agreement should describe what the vendor may do with conversation data. Name who controls it, who can access it, how long it is kept and whether it may train the vendor’s model or a third party’s model. Identify every model or hosting sub-processor. Record its location and the notice and approval process for changes. Set an export format, a deletion deadline and the evidence you will receive after exit. Ask whether backup copies follow the same schedule.

Map the data flow before approving the clauses: a customer asks about an order, the agent receives the message, may query your store and may send text to a model provider. Mark where personal information goes and who can read it. For the broader retailer obligations, see the Australian privacy guide. In this vendor check, the important point is that the contract and real data path must match.

For retailers covered by the Australian Privacy Principles, APP 8 generally requires reasonable steps before disclosing personal information to an overseas recipient to ensure that recipient does not breach the principles, subject to exceptions. The retailer can remain accountable for the recipient’s handling. Overseas server use is not automatically an APP 8 disclosure; the OAIC distinguishes disclosure from a use under the retailer’s effective control. Check the actual provider relationship with your privacy adviser.

APP 11 requires an APP entity to take reasonable steps to protect personal information it holds from misuse, interference, loss and unauthorised access, modification or disclosure. It also addresses destruction or de-identification when the information is no longer needed, subject to the stated exceptions. A vendor promise does not remove the retailer’s own need to check access, retention and deletion.

SaaS contract terms to negotiate

Read the software-as-a-service (SaaS) contract alongside the technical answers. Set an uptime commitment with a clear measurement period, exclusions, incident contact and remedy. Ask what happens if a service outage leaves the agent on a product page or mid-order. If a customer receives a wrong delivery promise or an unauthorised order action, the contract must explain responsibility, correction steps and any limit on liability. Agree who handles the customer while a dispute is resolved.

Ask how price changes at renewal, how much notice you receive and whether usage charges can rise with message or order volume. Require notice of material model, data-flow or sub-processor changes, with a way to assess the change before it affects live customers. Set termination notice, a usable export of conversations and configuration, assistance for migration and a deadline to revoke access and delete data. These are the contract answers that help you avoid AI vendor lock-in. Have your legal and privacy advisers review the final terms for your store and risk level.

Build a vendor scorecard

Set the weights before reading the vendor’s answers. The example below totals 100 points. Change the weights to match the access you plan to give the agent, then hold them fixed for every candidate. Give each criterion a score from 0 to 4: 0 means no credible answer, 1 a promise without evidence, 2 partial evidence, 3 evidence meeting your stated requirement and 4 evidence plus a successful pilot check. Multiply each score by its weight and divide by 4 to get weighted points. Record a reason and source beside every score.

CriterionWeightEvidence or question to record
Security20Which systems do the reports cover, and which serious findings remain open?
Privacy20Where does chat data go, who may use it and when is it deleted?
Answer quality20How often are pilot answers correct against your catalogue and policies?
Commerce actions15Which actions are allowed, confirmed and reversible?
Cost10What is the full first-year cost and renewal rule?
Support5Who responds to incidents and within what agreed time?
Exit10Can you export data and settings, receive exit help and revoke access?

Use the last column as a short vendor questionnaire. Send the same seven rows and request a named answer, evidence link or document and date for each. Add your store’s data-flow sketch, order permissions and pilot targets. The table is a reusable vendor risk assessment template: copy it into your working document or spreadsheet, add columns for score, evidence, risk owner and agreed fix, then retain the completed version with the contract. Agree the exit questions before signing, while you still have a choice of supplier.

A weighted total helps comparison, but it must not hide a serious risk. Put each open issue into a simple vendor risk assessment matrix. Rate likelihood and impact separately, using low, medium or high for each. For example, an agent permitted to submit an order without confirmation may have high impact even if the vendor says it is unlikely. Require an action limit or a written acceptance by the person authorised to own that risk. Use your agreed thresholds, not a score invented after the result.

Finish a one-page assessment report with the vendor and use case, total score, overall risk rating, evidence reviewed, open risks, agreed fixes, decision owner and next review date. Record which controls must be checked after a model change or incident. This turns a one-off procurement form into ongoing supplier monitoring.

Run a proof of concept before you sign

Run a time-boxed proof of concept or paid pilot with a defined start, end, cost and test owner. Use a real but bounded catalogue slice with current stock and policy information. Keep production customer data out until the data terms and access controls permit it. Agree in writing what happens to pilot conversations, exports and backups if you do not continue.

Build an AI agent evaluation set from real question types your store receives: a size or colour match, an out-of-stock item, a delivery exception and a request that the agent should hand to a person. Use anonymised examples where they contain personal information. Mark the expected source or action for each question, then score correctness, safe refusal or handover and time to correction. A chatbot testing method can help you prepare the cases. If the agent can act on a cart or order, test confirmation and permission limits in a safe pilot environment before live use.

Compare the pilot results with the targets and stop rules in the business case. Review wrong answers by type, not just the average score: one invented return promise may matter more than several correct colour suggestions. If the vendor changes the model, catalogue connection or action permissions, repeat the affected cases. Sign only when the written evidence meets your criteria. The contract and pilot results must meet them too.