12 Data Questions to Ask an AI Vendor

Most AI procurement reviews stall on the same handful of unanswered questions. Here they are, with what a defensible answer sounds like — useful whether you are the buyer asking or the vendor preparing.

Content and retention

1. Is prompt and completion content stored, and for how long? A good answer distinguishes content from metadata and gives a number. "We do not store your data" without that distinction is almost always inaccurate, because nobody bills or debugs without keeping something.

2. What operational metadata is kept? Expect: timestamps, token counts, model ids, latency, error codes, key identifiers. That set is normal and necessary. What matters is that it is enumerated rather than described as "some technical logs".

3. Is anything used to train models? The answer should be unambiguous and cover sub-processors. "We do not train on your data, and our upstream providers are configured not to either" is the shape you want, with the second half in writing.

4. Who internally can read content, and is that access logged? "Nobody" is rarely true — support engineers usually can, under a break-glass procedure. A vendor who admits that and describes the control is more credible than one who claims perfect isolation.

Sub-processors and geography

5. Who else touches the data? Model providers, cloud hosts, observability vendors, payment processors. Ask for the list, and ask how you will be told when it changes.

6. Where is data processed? Region matters for regulated workloads, and the honest answer is often "wherever the provider serves the model", which may not be the same region as the vendor's own infrastructure.

7. What happens on termination? Deletion timelines for content, metadata, and backups — backups are the part usually forgotten, and the part with the longest tail.

Security posture

8. What certifications are actually held? Note the word held. "SOC 2 in progress" and "SOC 2 certified" are different statements, and a vendor blurring them is telling you how they handle other claims. A young company with none, stated plainly, is more trustworthy than one implying otherwise.

9. How are credentials stored and access controlled? Especially relevant for anything holding your provider keys — see key management for what to probe.

10. Is there a disclosure channel and a track record? A published security contact and a stated response window. Nothing exotic; its absence is the signal.

Contracts and continuity

11. What can be signed? MSA, DPA, and whether an order form can override standard terms. Also: which legal entity, in which jurisdiction, and can it invoice in your currency.

12. What is the exit path? Export of usage history, notice period, and whether the integration is standard enough that leaving is a configuration change. A gateway speaking a standard API is a base-URL change to replace; a proprietary SDK is a migration project.

Two answers that should stop the process

The first is a specific claim that cannot be evidenced — a certification without a report, an uptime figure without a status page, a "zero retention" claim that dissolves when you ask about metadata. If a checkable claim turns out to be decorative, treat the uncheckable ones accordingly.

The second is a refusal to put an answer in writing. Everything above should survive being pasted into a contract schedule. A vendor comfortable saying something on a call but not in a document has told you which version they intend to be bound by.

For our own answers, the security page and privacy policy are written to be quotable, and where we hold nothing we say so.

Sending us this questionnaire is welcome. Runix Gateway is in early access — tell us what you are building and we reply within one business day.