10 Questions to Ask Before Connecting AI to Your Business Data

September 24, 2026

10 Questions to Ask Before Connecting AI Banner

Before connecting any AI feature, assistant, or integration to business data, put three facts in writing:

  • Who covers the processing, including whose agreement governs the AI path.
  • What is retained, including prompts, uploads, files, and knowledge bases.
  • What is audited, including whether AI access to records is logged and exportable.

These are not abstract governance questions. They determine what happens when customer records, financials, or regulated data start flowing through a model. The good news is that the answers are knowable: major model providers publish data terms, and the differences between tiers, endpoints, and defaults are documented in their own words.

The risk is that those differences are often buried where buyers rarely look. The same vendor can treat your data differently depending on the tier, endpoint, or entry point you use. The ten questions below surface those differences before connection day. For each one, this guide explains why it matters, what a good answer sounds like, and what a bad answer sounds like.

The AI Data Checklist: 10 Questions Every Business Should Ask

These questions focus on the three areas that matter most when evaluating AI access to business data: contractual coverage, data handling, and accountability. A vendor that can answer them clearly and in writing is easier to evaluate than one that relies on marketing claims.

1. Whose API key does the AI run on?

There are two common architectures, and they assign responsibility differently. If the AI runs on your own API key, you hold the relationship with the model provider: you inherit its data terms directly, choose the tier, and remain responsible for reading and negotiating what you signed. If the AI runs as the platform’s covered processing, the platform holds the agreement with the model provider and passes its terms down to you. The first model gives you control and direct accountability. The second can simplify procurement, but only if you verify what the platform actually agreed to on your behalf.

A good answer sounds like: “The feature runs on [provider]’s API under our agreement with them; here is the document describing the tier, the retention terms, and what flows down to you.”

A bad answer sounds like:We use GPT-5.” A model name is not an answer. It says nothing about whose agreement governs the data, which tier it runs on, or who is accountable when terms change.

2. Is there a business agreement covering the AI processing path, and what does it scope?

Where regulated data is involved, the question is not whether the model is good but whether a signed agreement, a BAA where HIPAA applies, covers the specific path your data takes. Model providers may sign agreements covering specific API services for eligible customers, but coverage is per agreement and per service, never blanket, and consumer chat products are not covered by it. “The model provider supports HIPAA workloads” is a statement about what is possible, not about what your vendor signed. Ask which entity signed, which services and endpoints the agreement scopes, and whether the AI feature you are evaluating stays inside that scope. Our companion piece on what a BAA actually covers breaks down what these agreements do and, just as important, what they leave on your side of the table.

A good answer sounds like: A signed agreement naming the covered services, plus a plain statement of which AI paths fall outside it.

A bad answer sounds like: A link to the model provider’s trust page, offered as if it were a contract.

3. What is the data retention posture at the model provider?

Data retention usually falls into three buckets:

  • Standard retention: Logs are kept for a defined period, often for abuse monitoring.
  • Reduced or opt-out retention: Retention is shortened or limited under specific terms.
  • Zero data retention (ZDR): Request data is not retained, but this typically requires the provider’s prior approval and applies only to eligible endpoints.

For example, OpenAI’s API documentation describes default abuse-monitoring retention of up to 30 days on most endpoints, with provisioned data controls, including ZDR, available for approved use cases on eligible endpoints (OpenAI, “Your data,” accessed July 2026).

Whatever posture applies, get it in writing and make sure the document is dated. Terms can change over time, and a dated record gives you a baseline for comparison. That matters because the FTC warned in February 2024 that “it may be unfair or deceptive for a company to adopt more permissive data practices” and “only inform consumers of this change through a surreptitious, retroactive amendment to its terms of service or privacy policy” (FTC; the elided passage gives AI training as one example).

A good answer sounds like: “Retention on your path is X days for abuse monitoring, or zero under our approved ZDR arrangement; here it is in the agreement.”

A bad answer sounds like: “We don’t store anything,” said out loud, matching nothing in any written term.

4. Do file and knowledge uploads follow the same rules as prompts?

This is the question almost nobody asks, and it is where retention promises quietly fail. Uploads and stateful features often persist under different rules than prompts. OpenAI’s data controls documentation draws the structural line: stateless endpoints such as chat completions, responses, and embeddings are eligible for zero data retention, while stateful endpoints, including conversations, assistants, threads, vector stores, and files, retain application-state data regardless. And the rules attached to each side move.

OpenAI’s earlier public BAA form (November 2024) scoped eligible API use to its “Zero Retention API,” which kept those stateful endpoints out of PHI processing; a new HIPAA guide posted July 9, 2026 replaced that gate with a family of four “Modified Retention” controls (Modified Abuse Monitoring, Zero Data Retention, Safety Retention, or Eyes Off), and OpenAI’s HIPAA-eligible endpoint list, posted the same day, now includes conversations, threads, vector stores, and files, which “can be used for processing PHI, even if data is retained” once an account is provisioned and a BAA is executed (OpenAI HIPAA Guide, July 9, 2026; endpoint list, same date). The endpoints may be the same, but the governing rules changed materially, and the change can be tied to a specific date.

Microsoft documents the same structural split for Azure OpenAI: the models themselves are stateless, but stateful features such as file storage, vector stores, and assistants threads store data at rest until the customer deletes it (Microsoft). Google’s paid Gemini API tier is a useful contrast because it names uploads inside the protection: “Google doesn’t use your prompts (including associated system instructions, cached content, and files such as images, videos, or documents) or responses to improve our products” (Gemini API Additional Terms). The lesson is double: retention terms are per-feature, and they move within weeks, so get each feature’s terms in writing, dated.

A good answer sounds like: A separate, explicit, dated retention statement for files and knowledge bases, plus a policy on whether regulated data may ever enter uploads.

A bad answer sounds like: One blanket retention claim with no mention of uploads at all.

5. Is your data used for model training, and under which tier?

Whether your data may be used for training is determined by the tier, not the vendor alone, and the same company routinely runs opposite defaults on different tiers. OpenAI states: “As of March 1, 2023, data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in to share data with us),” while consumer ChatGPT includes an on-by-default “Improve the model for everyone” setting. Anthropic’s Commercial Terms state that “Anthropic may not train models on Customer Content from Services,” while its August 2025 consumer policy update allows Free, Pro, and Max chats to be used for training unless the user opts out, with retention extending to five years for those who allow it (Anthropic Commercial Terms; consumer update).

Google’s unpaid Gemini API tier states that Google “uses the content you submit to the Services and any generated responses to provide, improve, and develop Google products and services,” that human reviewers may read your input and output, and then states the implication directly in its terms: “Do not submit sensitive, confidential, or personal information to the Unpaid Services.” That sentence, from the vendor itself, is the entire argument for this question.

A good answer sounds like: “Your data runs on [tier], and here is that tier’s training clause in the contract.”

A bad answer sounds like: “We never train on your data,” with no tier named and no clause cited.

6. Are AI connections on by default or off by default, and who can enable them?

Defaults often determine outcomes in most organizations, because most settings are never touched. The documented examples are striking. Figma’s content-training settings ship opposite defaults by plan: “By default, content training is turned on for Starter teams” and Professional teams, and turned off on Organization and Enterprise plans (Figma). Slack states plainly that “We do not develop generative AI models using Customer Data,” but for its non-generative machine-learning models (search ranking, channel and emoji recommendations), customer data contributes by default, and opting out requires the workspace owner to email Slack; there is no in-product toggle (Slack privacy principles). Ask what is on the day you sign, and who in your organization has the authority to flip it. This question belongs inside a broader policy; see our vibe coding governance checklist for where it fits.

A good answer sounds like: “AI features are off until an admin enables them, per workspace, and here is the settings documentation.”

A bad answer sounds like:Users can manage that themselves,” which means it is on, everywhere, now.

7. Is there an audit trail of AI access to records, and can you export it?

When an AI feature reads a record, it is accessing business data and should leave the same kind of trail a person would. The log should show who or what accessed which record, when it happened, and on whose behalf. Without that record, you cannot answer an auditor, investigate an incident, or verify the vendor’s own claims about what the AI touched. Exportability matters too: a log you cannot pull into your own review process is a screenshot, not evidence.

A good answer sounds like: “AI access appears in the same exportable audit log as user access, with each request attributable to the user, role, or service that triggered it.”

A bad answer sounds like: “We can look into usage on our side if something comes up.” That is vendor support, not an audit trail.

8. Does the AI inherit your access model, or does it see everything?

Your application presumably enforces roles and record-level security: a regional manager sees her region, not the company. The question is whether the AI operates inside that same model, seeing only what the requesting user could see, or whether it queries the data layer with broad privileges and becomes a bypass around every permission you configured. An AI layer with unrestricted access turns your carefully scoped app into an open-book interview: anyone who can type a question can retrieve anything the model can reach.

A good answer sounds like: “The AI executes in the context of the signed-in user and inherits roles and record-level permissions.”

A bad answer sounds like: “The AI uses a service account with read access to the database.”

9. Where exactly is the boundary between covered and customer-controlled AI paths?

Almost every platform has two kinds of AI paths: ones the vendor covers contractually, and ones you can wire up yourself (your own keys, your own integrations) that sit explicitly outside the vendor’s coverage. Neither is wrong. What matters is that the boundary is written down, because the failure mode is discovering after an incident that the path your team used was on the uncovered side.

The strongest vendors document that boundary publicly, and the line moves. Airtable’s Health Information Datasheet stated in March 2026 that “Airtable AI is not currently available for customers who require a Health Information Exhibit”; a July 23, 2026 update to the same document now says “Airtable AI is available to customers with an executed Health Information Exhibit… including the Supplemental AI Terms for Health Information” (Airtable Health Information Datasheet, accessed July 2026). Same vendor, same page, different answer, four months apart. That is not a criticism; it is the whole argument for getting the boundary in writing with a date, because a vendor that states it plainly has answered the question, and the answer you hold is only as current as its date.

A good answer sounds like: A written statement of which AI features are inside the vendor’s agreements and which customer-controlled paths are outside them.

A bad answer sounds like: “Everything is covered,” with no list of what “everything” contains.

Diagram of two AI paths from one app and its business data: a platform-covered path inside platform coverage, where the vendor holds the agreement and its terms flow down, and a customer-controlled path on your own keys outside coverage, divided by a boundary line asking whether the boundary is written down.

10. Can you turn it off, and at what granularity?

An all-or-nothing kill switch is a blunt instrument. Real organizations need selective control: AI enabled for the sales app but not the HR app, for general business tables but not the table holding sensitive records, for analysts but not contractors. Granular disablement also provides flexibility if provider terms change over time. If a provider’s posture shifts at renewal, you want to narrow exposure without shutting down every AI workflow the business now depends on.

A good answer sounds like: “Per app, per table or data source, and per role, controlled by your admins.”

A bad answer sounds like: “You can request that AI be disabled for your account by contacting support.”

Red Flags: When to Stop the Evaluation

Any one of these is a reason to pause; two or more is a pattern.

  • Silence on retention. If retention is not stated in a document you can keep, assume standard retention and assume you will not be notified when it changes.
  • Training ambiguity. “We don’t train on your data” without a named tier and a citable clause is not sufficient due diligence. All major providers’ protections are tier-scoped; an answer that ignores tiers has not read its own paperwork.
  • No uploads answer. A vendor that quotes prompt retention but goes quiet on files and knowledge bases either has not checked or does not want to say. The providers’ own documentation treats these differently; your vendor must too.
  • No audit trail. If AI access to records is unlogged, every other assurance is unverifiable by construction.
  • Everything on by default. Defaults that favor the vendor’s data position, combined with opt-outs that require emails or support tickets rather than settings, tell you how the rest of the relationship will go.
  • Compliance by proximity. Pointing at a model provider’s certifications as if they transferred automatically. Agreements cover named services for named parties; nothing transfers by association.

Platforms That Answer in Writing

Every question above has either a written answer, or it does not. That is the test. The strongest signal a platform can send is to publish those answers before buyers have to ask: which provider and tier its AI runs on, which agreement covers the path, what is retained, what is logged, and where the covered boundary ends.

Caspio is one example of a platform that answers these questions in writing. Its AI features run on OpenAI’s API under a signed BAA, operate inside the platform’s existing access model and audit logging, and give administrators control over where AI is enabled. At the feature level, Caspio’s in-app AI Agent makes question 8 concrete: it executes in the signed-in user’s role, inherits record level security, and applies field-level security at design time. On question 4, the uploads question, there is now a written answer too: Caspio’s updated BAA covers file and knowledge-base uploads for HIPAA-covered accounts, which is the kind of separate, dated uploads answer this guide says to demand. If you are evaluating this category against regulated workloads, our guide to AI app builders for regulated industries applies the same written-evidence standard to the whole vendor field.

Run the 10 questions against any vendor tomorrow. The ones worth connecting to your data will hand you documents. The rest will hand you adjectives.

Frequently Asked Questions

Is it safe to connect AI to company data?

It can be, if three things are established in writing first: who covers the processing (whose agreement governs the AI path), what is retained (for prompts and, separately, for uploads), and what is audited (logged, attributable AI access to records). Business and API tiers of major model providers often offer no-training defaults and defined retention policies, while consumer AI products generally follow different rules.

Does AI train on my data?

It depends on the tier, not the vendor. OpenAI and Anthropic both exclude API and commercial-tier customer content from training by default while running different rules on consumer tiers, and Google’s unpaid Gemini API tier uses submitted content to improve its products while the paid tier does not. Never accept a training answer that does not name the tier and cite the clause.

What is zero data retention (ZDR) and do I need it?

ZDR is an arrangement, typically requiring the model provider’s prior approval, under which eligible API endpoints retain no request data, not even short-term abuse-monitoring logs. It is not universal. Many stateful endpoints and upload-based features fall outside ZDR eligibility. And it is no longer the only gate. As of July 9, 2026, OpenAI’s HIPAA terms accept any of four Modified Retention controls, ZDR among them. Ask which control applies to your path, and get it in writing, dated.

Should we use our own API key or the vendor's built-in AI?

Your own key gives you the direct provider relationship. You pick the tier, inherit its terms, and carry the responsibility for reading and maintaining them. Vendor-covered AI is simpler and puts the agreement burden on the platform, but you must verify what the platform signed, which services it scopes, and whether its terms flow down to you in writing. Either can pass review; unexamined versions of both fail it.

What should an AI vendor due diligence checklist include?

At minimum, it should document several things: whose API key and agreement govern the processing path; whether a BAA or equivalent agreement covers that path where regulated data applies; the retention posture, in writing and dated; separate rules for uploads versus prompts; training use by tier; on/off defaults and who controls them; an exportable audit trail of AI access; whether the AI inherits your access model; the boundary between covered and customer-controlled paths; and how granularly AI can be disabled.

Call to Action Block Call to Action Block

Recommended Articles

Caspio AI Chat Agent banner

Add a Role-Aware AI Chat Agent to Your Caspio App

READ STORY
Lovable Alternatives Banner

Lovable Alternatives for Business and Regulated Apps (2026)

READ STORY

Vibe Coding vs. Low-Code: Key Differences and Where Caspio Fits in 2026

READ STORY

Vibe Coding Governance: The IT Leader’s 2026 Checklist

READ STORY

Code Artifact vs. Governed Platform: The Two Architectures

READ STORY
Shadow AI Apps Banner

Shadow AI Apps Are the New Shadow IT: A 2026 Guide

READ STORY

Secure Alternatives to Vibe Coding for Business Apps (2026)

READ STORY
AI App Builders for Regulated Industries Banner

AI App Builders for Regulated Industries: 2026 Buyer's Guide

READ STORY
AI & No Code Banner

AI and No-Code: Generative AI in App Development

READ STORY

Per-User Pricing vs Flat Rate: The Unlimited Users Math

READ STORY

Rebuild or Retrofit After a Failed Security Review (2026)

READ STORY
Hidden Cost of Free AI App Builders 2026 Banner

The Hidden Costs of Free AI App Builders (2026)

READ STORY
Subscribe for More Updates