SiteGPTStart free trial

12 Questions to Ask a Chatbot Vendor Before You Sign a BAA

A twelve-question due diligence set for healthcare teams evaluating a chatbot vendor, grouped into contract, data, product, and incident, with SiteGPT's own answers given in full.

Sai Dheeraj

SiteGPT Team

Questions Before You Sign a BAA

SiteGPTBest AI chatbot for customer service

Every chatbot vendor selling into healthcare will tell you they are HIPAA compliant. The phrase is unregulated, no one audits it, and it is printed on the pricing pages of products that will not sign a Business Associate Agreement at any price.

A few terms first, because the rest of this depends on them. HIPAA is the Health Insurance Portability and Accountability Act, the US law governing how patient information is handled, and PHI, or protected health information, is any health detail traceable back to a person. A business associate, often shortened to BA, is an outside vendor that handles PHI on your behalf; a BAA, or Business Associate Agreement, is the contract that permits it. HHS, the Department of Health and Human Services, writes the guidance and enforces the rules.

What follows is a twelve-question screen you can send to any vendor as written. SiteGPT's own answers are included under each one, the least flattering ones included, because a question set where the publisher scores perfectly is a sales sheet.

iShort answer

Twelve questions, grouped into contract, data, product, and incident, that separate a chatbot vendor who can support a HIPAA deployment from one who has written "HIPAA compliant" on a page. Three answers should end the evaluation: refusing to sign a BAA while still marketing to healthcare, being unable to name the subprocessors that touch conversation data, and claiming nothing changes in a covered workspace. SiteGPT answers all twelve below, including no EU data residency, BAA on the Enterprise plan only, and no automated PHI detection anywhere in the product.

The 12-question BAA screen

Four bands, three questions each. Work them in order; the contract questions decide whether the rest are worth asking.

1

The contract (Q1-Q3)

Will they sign, on which plan, does the coverage flow down to every subprocessor including the AI model provider, and can you see the SOC 2 Type II report and subprocessor list.

2

The data (Q4-Q6)

Which countries the conversations are stored and processed in, how long content is retained and whether that is a contract term, and whether anything is used to train models.

3

The product (Q7-Q9)

Whether leads are governed differently from transcripts, which features a covered workspace disables, and where a conversation goes when it is escalated to a human.

4

The exit and the incident (Q10-Q12)

What happens to PHI at cancellation, how many days until a breach is reported and counted from what, and who inside the vendor can reach your workspace.

Key takeaways

The question that settles it fastestWhat you are actually testing
Will you sign a BAA, and on which plan?Whether the marketing claim survives contact with procurement. Plan gating surfaces here or three weeks later.
Name every subprocessor that can touch PHI.Whether the vendor knows its own chain. Flow-down makes the AI provider and the host part of the covered set.
Which regions store and process the data?Country names, not adjectives. This is where residency requirements are met or missed.
Is retention a contract term or a settings toggle?Whether an administrator can quietly change what you signed for.
What is disabled in a covered workspace?Whether a real HIPAA mode exists. 'Nothing changes' means it does not.
Which SiteGPT plan?The Enterprise plan, at custom pricing. Not Starter, Growth, or Scale.

How to use this list

Send it before the demo, not after. Vendors who cannot answer Q1 and Q2 will consume three weeks of evaluation time before that becomes clear, and the answers to those two are a single email away.

Ask for written answers. A sales engineer saying "yes, we're fully compliant" on a call is not something your privacy officer can file, and the difference between a verbal yes and a written one is precisely the difference this list is designed to expose.

Judge specificity, not confidence. Across all twelve questions, the useful signal is whether the answer contains a number, a country, a named company, or a clause reference. Answers made of adjectives are answers the vendor has not had to defend before.

The contract (Q1-Q3)

These three establish whether there is anything to evaluate. A vendor that clears all three has a real compliance program; a vendor that stumbles on any of them has a marketing claim.

Q1. Will you sign a BAA, and on which plan?

Why it matters. HIPAA permits you to disclose PHI to a business associate only once you have satisfactory assurances in a written agreement, so a vendor who will not sign cannot be used with patient data regardless of how good the product is. The plan question is the follow-up most teams forget, and it is where budget assumptions break.

A good answer names the plan and the price model. A weak answer is "yes, we support HIPAA", which is not the same sentence.

SiteGPT's answer. Yes, on the Enterprise plan only, at custom pricing. BAAs are not available on Starter, Growth, or Scale, and a self-serve HIPAA tier is planned rather than available. The standard agreement is published in full at sitegpt.ai/legal/baa, so counsel can read the text before anyone books a call, and it is a standard document rather than a fixed one: retention windows, breach notice timelines, notice contacts, and scoping are commonly adjusted during onboarding.

Q2. Does the BAA cover every subprocessor that can touch PHI, including the LLM provider?

Why it matters. A chatbot is never one company. There is a hosting provider underneath it, a vector database beside it, and an AI model provider generating the answers, and HIPAA covers them through flow-down: a business associate must hold a BAA with any subcontractor before disclosing PHI to it, and every downstream subcontractor is a business associate in its own right.

A good answer names companies. You are not signing those contracts yourself, so what you need is written confirmation that the chain is complete and a list you can check it against.

SiteGPT's answer. Every vendor in the PHI chain operates under its own BAA, and SiteGPT countersigns only once that coverage is in place. The full list is published at sitegpt.ai/legal/subprocessors rather than supplied on request, so you can verify the AI layer yourself instead of taking a summary of it.

Q3. Can we see the SOC 2 Type II report and the current subprocessor list?

Why it matters. SOC 2 Type II tests whether controls operated over a period, not whether they existed on one day, which is why it is the report an auditor asks for. Pairing it with the subprocessor list is deliberate: one tells you the controls work, the other tells you who is inside them.

A good answer produces both. A Type I report offered in place of a Type II, or a subprocessor list available "under NDA after signature", are both answers worth noting.

SiteGPT's answer. A SOC 2 Type II examination covering the security, availability, and confidentiality trust service criteria was completed with zero exceptions noted across all tested controls. The report is available through the trust portal at trust.sitegpt.ai, and the subprocessor list is public.

The data (Q4-Q6)

The contract band tells you whether a vendor will commit. This band tells you what they are committing to.

Q4. Where is conversation data stored and processed? Name the regions.

Why it matters. "Secure cloud infrastructure" is not an answer to this question. Regions decide cross-border transfer analysis, and for organizations with state-level or EU requirements layered on top of HIPAA, they decide whether the vendor is usable at all.

A good answer is a list of countries. Ask separately about storage and processing, because inference often happens somewhere the database does not.

SiteGPT's answer, and it is the weakest one here. There is no regional hosting option and no data residency commitment. The published subprocessor list names no EU-based subprocessor: Convex and Pinecone run in the United States on AWS, OpenAI and Cohere are United States, Cloudflare is listed as Global, and Paddle is United Kingdom. If your requirement is EU-only storage and processing, SiteGPT does not meet it today, and you should confirm that against the list rather than against this paragraph.

Q5. What is the default retention for conversation content, and can we change it?

Why it matters. HIPAA does not set a retention period for conversation content, which means the vendor's default is the operative rule until you negotiate it. The more revealing half of the question is the second half: a number written into the order form and a number an administrator can change in settings are different guarantees.

A good answer gives both the default and the mechanism.

SiteGPT's answer. Conversation content is redacted seven days after a conversation's last activity by default, and the window is configurable on the order form, which makes it a contract term rather than a dashboard toggle. Redaction replaces the message text while preserving conversation counts and analytics aggregates, so reporting survives the deletion.

Q6. Is customer data ever used to train or fine-tune models?

Why it matters. This is the question with the widest gap between policy and enforcement. The vendor may not train on your data while the AI provider behind it retains prompts for abuse monitoring, and both statements can be made truthfully at once.

A good answer covers the whole chain and says what enforces it. Ask what is technical and what is contractual.

SiteGPT's answer. Conversations, training data, and customer interactions are never used to train or fine-tune any AI model, and this is a commitment in the BAA rather than only a policy page. It is backed at the AI layer by zero-data-retention endpoints, including with OpenAI, which are named on the public subprocessor list.

0

EU-based subprocessors on SiteGPT's published list, and zero exceptions noted in the SOC 2 Type II examination. Two numbers from the same source pattern: publish the list, let the buyer check it, accept the answer that follows.

Source: SiteGPT subprocessor list and security page

The product (Q7-Q9)

Contract and data questions get asked. Product questions usually do not, which is why they are where evaluations go wrong after signature.

Q7. Are leads and captured contacts treated differently from conversations?

Why it matters. Retention policies get written for transcripts. Lead records, the name plus email plus "what brings you in today" that the widget captured, sit in a different table and frequently outlive every retention rule the contract describes.

A good answer acknowledges the distinction rather than being surprised by it.

SiteGPT's answer. Yes, and they are governed differently on purpose. Conversation content is redacted on the retention window; leads persist until you delete them, because a lead is a deliberate contact submission rather than an incidental transcript. That puts the deletion decision with your organization, which can export leads to CSV and delete them from the dashboard against its own retention obligations. It also means a lead table is somewhere patient data can accumulate quietly, so it belongs in your periodic review.

Q8. Which product features are disabled in a covered workspace?

Why it matters. Every honest HIPAA configuration is narrower than the standard product, because the standard product is optimized for convenience and a covered one is optimized for a known data path. A vendor who says nothing changes is describing a product where PHI can reach anywhere.

A good answer is a specific list, paired with the alternative for each removal.

SiteGPT's answer. Cloud drive sources (Notion, Google Drive, Dropbox, OneDrive, Box, GitHub) are off, and chat integrations (Zendesk, Slack, Crisp) are off. Sources already connected to an account are switched off when HIPAA is enabled, and lead form templates are hidden so fields must be chosen deliberately. Content instead arrives by upload, paste, or a crawl of your own site, and chats run in SiteGPT's own widget. The "Responses are AI-generated" disclosure stays pinned even with white-label branding.

Q9. How does human escalation work when a transcript contains PHI?

Why it matters. Escalation means moving a conversation somewhere a person will read it, and that destination is usually a different product. If the handoff target sits outside the covered chain, the compliant chatbot has become a delivery mechanism for an uncovered disclosure.

A good answer names where the transcript lands and confirms that location is covered.

SiteGPT's answer. Chat history review, human takeover, and chat modes work normally in a covered workspace, and the handoff happens inside SiteGPT rather than by forwarding the conversation into Zendesk, Slack, or Crisp, which are off. That is a narrower workflow than the standard product's, and it is narrower on purpose: it keeps the transcript inside the covered boundary instead of extending the boundary to another vendor.

The exit and the incident (Q10-Q12)

These three describe what happens on the worst day of the relationship and the last day of it. They are the least likely to be answered on a website.

Q10. What happens to PHI when we cancel?

Why it matters. Return or destruction of PHI at termination is a required BAA element under 45 CFR 164.504(e), so every vendor's contract will contain the clause. The clause is not the answer. The mechanism and the timeline are.

A good answer describes who does what, and by when.

SiteGPT's answer. Return or destruction of PHI at termination is committed in the BAA. Day to day, your organization can independently delete individual conversations, trained content, leads, or the entire chatbot from the dashboard, so the exit does not depend on filing a request and waiting.

Q11. How fast do you report a breach, and to whom?

Why it matters. You owe individual notice without unreasonable delay and no later than 60 days after discovery, and if more than 500 residents of a state are affected, media outlets and the Secretary of HHS as well. A vendor clock that consumes most of your window is a contract problem, and "promptly" is not a clock.

A good answer is a number of days, counted from a defined event.

SiteGPT's answer. Within five business days of discovery, plus security incident reporting, both written into the BAA. Breach notice timelines are among the provisions commonly adjusted during onboarding, so if five days does not leave your team enough room, that is a negotiation rather than a dead end.

Q12. Who inside your company can access our workspace, and is that access logged?

Why it matters. External controls get documented because buyers ask about them. Internal access is the control most likely to be undocumented publicly and most likely to be examined in an audit or after an incident.

A good answer separates customer-side permissions from vendor-side access, and states whether vendor access is logged and reviewable.

SiteGPT's answer, in two halves. On your side, role-based permissions control which team members can reach chatbot settings, training data, and conversation history. On SiteGPT's side, the specifics of internal access and its logging are not detailed on the public security page. Those controls are within the scope of the SOC 2 Type II examination, so the report available through trust.sitegpt.ai is the right artifact to request, and this is a question to get answered in writing during onboarding rather than to consider answered by this page.

The three answers that should end the evaluation

Most answers on this list are comparison signals. Three are not.

ScenarioBest pickWhy
"We don't sign BAAs, but we're HIPAA compliant."Walk awayThe two halves contradict each other. Without a BAA you cannot lawfully disclose PHI to that vendor, so the compliance claim describes a configuration you are not permitted to use.
"We can't share which subprocessors touch conversation data."Walk awayFlow-down requires the vendor to have a BAA with every subcontractor handling PHI. A vendor who will not name the chain is telling you it cannot demonstrate the chain is covered.
"Nothing changes in HIPAA mode, all features work the same."Walk awayA covered configuration constrains where PHI can travel, which always removes something. This answer means either no covered mode exists or nobody has mapped the data paths.

Two more answers are yellow rather than red. "Retention is configurable in settings" means the guarantee lives where an administrator can change it, so ask for the number in the order form. And "we're SOC 2 certified" without a Type II report attached is worth one follow-up, because Type I and Type II are different examinations and the shorthand hides which one was done.

The full question set, to copy

Send as-is. It is written to be answerable in a single reply.

Before we can proceed to a BAA, we need written answers to the following.

CONTRACT
1. Will you sign a BAA, and on which plan or tier?
2. Does your BAA cover every subprocessor that can touch PHI,
   including the LLM/AI model provider? Please name them.
3. Can you provide your SOC 2 Type II report and current
   subprocessor list?

DATA
4. In which countries is conversation data stored, and in which
   is it processed? Please answer both separately.
5. What is the default retention period for conversation content?
   Is that a contract term or an in-product setting?
6. Is customer data ever used to train or fine-tune models, by you
   or by any AI provider you use? What enforces that technically?

PRODUCT
7. Are leads and captured contacts subject to the same retention
   rules as conversation transcripts? If not, what governs them?
8. Which product features or integrations are disabled or altered
   in a HIPAA-enabled workspace?
9. When a conversation is escalated to a human, where does the
   transcript go, and is that destination inside the BAA scope?

EXIT AND INCIDENT
10. What is the process and timeline for return or destruction of
    PHI when we terminate?
11. Within how many days of discovery do you report a breach, and
    to whom? Is that counted in business days or calendar days?
12. Which of your personnel can access our workspace and its data,
    under what circumstances, and is that access logged and
    reviewable by us?

Where SiteGPT scores badly on its own list

Publishing a due diligence set and then passing it perfectly would tell you the set was written backwards. Here is the honest split.

Pros

  • The BAA is published in full at sitegpt.ai/legal/baa, readable by counsel before any sales conversation
  • It is a standard document rather than a fixed one: retention windows, breach timelines, notice contacts, and scoping are commonly adjusted at onboarding
  • Breach reporting within five business days of discovery, a number rather than 'promptly'
  • Subcontractor flow-down, with SiteGPT countersigning only once the full PHI vendor chain is covered
  • PHI never used to train AI models, backed by zero-data-retention endpoints at the AI layer
  • SOC 2 Type II completed with zero exceptions, and a public subprocessor list you can check without asking
  • Conversation retention is a contract term on the order form, not a toggle an administrator can change

Cons

  • No EU data residency and no regional hosting option; the subprocessor list names zero EU-based subprocessors
  • The BAA is Enterprise plan only, at custom pricing; Starter, Growth, and Scale do not include it
  • A self-serve HIPAA tier is planned and is not available today
  • Cloud drive sources and chat integrations are off by default in a covered workspace, and switch on only for vendors you hold your own BAA with
  • Sources already connected to an account are switched off when HIPAA is enabled
  • Nothing scans customer-supplied content for PHI; keeping training content, bot instructions, and lead fields clean is your review step
  • Internal access controls and their logging are not detailed publicly; that answer comes from the SOC 2 report and onboarding

The pattern worth noticing across both columns is that the strengths are all things you can verify without talking to anyone, and the weaknesses are all things a vendor could have declined to volunteer. That is the same test to run on every other vendor's answers.

Best forUS healthcare compliance and procurement teams running a structured vendor evaluation before a chatbot handles any patient data.

If you are earlier than vendor selection and still establishing whether a BAA is required at all, start with do you need a BAA for your website chatbot. For the full requirements picture, including safeguards and the tradeoffs of a covered workspace, see what makes a chatbot HIPAA compliant. If your knowledge base is the open question, keeping PHI out of your chatbot's training data covers that specifically. SiteGPT's own program status lives at the HIPAA program page and the day-to-day behavior of a covered account at the HIPAA workspace documentation.

Frequently asked questions

What should you ask a chatbot vendor before signing a HIPAA BAA? Work through four areas rather than a single list. The contract: will they sign a BAA, on which plan, does it cover every subprocessor including the AI model provider, and can you see the SOC 2 Type II report. The data: where it is stored and processed, how long conversation content is retained, and whether anything is used to train models. The product: how leads differ from conversations, which features are disabled in a covered workspace, and how human escalation handles a transcript containing PHI. The incident and exit: what happens to PHI at cancellation, how fast a breach is reported, and who inside the vendor can access your workspace.

Does a HIPAA BAA need to cover the AI model provider? Yes, though not through a contract you sign yourself. HIPAA handles this through flow-down: a business associate must have a BAA with any subcontractor before disclosing PHI to it, and every downstream subcontractor is a business associate in its own right. So the AI model provider is your chatbot vendor's responsibility to cover, and your job is to confirm in writing that the chain is complete. Ask for the subprocessor list and check that the AI providers on it are named.

What answers should disqualify a chatbot vendor? Three. First, refusing to sign a BAA at all while still marketing to healthcare, because the marketing claim and the contract position cannot both be true. Second, being unable to name the subprocessors that can touch conversation data, since a vendor who cannot list the chain cannot have covered it. Third, claiming that nothing changes in a HIPAA workspace, because every honest covered configuration is narrower than the standard product and a vendor who says otherwise has either not built one or is not describing it accurately.

Does SiteGPT offer EU data residency for chatbot conversations? No. The published subprocessor list names no EU-based subprocessor: Convex and Pinecone run in the United States on AWS, OpenAI and Cohere are United States, Cloudflare is listed as Global, and Paddle is United Kingdom. There is no regional hosting option and no residency commitment, so a requirement for EU-only storage and processing is not one SiteGPT meets today. Teams with that requirement should confirm it against the subprocessor list themselves rather than take any vendor's summary of it.

How long does SiteGPT keep chatbot conversations under HIPAA? Conversation content is redacted seven days after a conversation's last activity by default, and the window is configurable on the order form, which makes it a contract term rather than a dashboard setting. Redaction replaces the message text while preserving conversation counts and analytics aggregates. Leads are treated separately: they persist until deleted, because they are deliberate contact submissions, and organizations manage that deletion themselves against their own retention obligations.

Which features are disabled in a HIPAA-enabled SiteGPT workspace? Cloud drive sources such as Notion, Google Drive, Dropbox, OneDrive, Box, and GitHub are off, as are chat integrations such as Zendesk, Slack, and Crisp. Sources already connected to an account are switched off when HIPAA is enabled, and lead form templates are hidden so fields have to be chosen deliberately. These are defaults rather than prohibitions: a specific source can be enabled for a vendor you hold your own BAA with. Content otherwise arrives by upload, paste, or a crawl of your own site, and chats run in SiteGPT's own widget.

Which SiteGPT plan includes a HIPAA BAA? The Enterprise plan only, at custom pricing. BAAs are not available on Starter, Growth, or Scale, and a self-serve HIPAA tier is planned rather than available. The standard BAA is published in full at sitegpt.ai/legal/baa so counsel can review the text before any sales conversation, and it is a standard document rather than a fixed one: retention windows, breach notice timelines, notice contacts, and scoping are commonly adjusted during onboarding.

Does SiteGPT scan for PHI in training content? No, and no vendor should be assumed to. Nothing scans customer-supplied content for patient data, so keeping it out of training content, bot instructions, and lead fields is your own review step. This matters because training content is what the bot retrieves from, which means anyone who can chat with the bot can reach it indirectly. If patient data does reach a lead record it is recoverable rather than catastrophic, since leads are never deleted on a schedule and can be found and removed.

Sources

  • HHS guidance on business associates for the business associate definition, the written agreement requirement, and subcontractor flow-down, read 8 August 2026
  • HHS Breach Notification Rule for the 60-day individual notice deadline and media notice above 500 residents of a state
  • SiteGPT HIPAA program page for plan gating and program status, verified 8 August 2026
  • SiteGPT standard BAA for the five business day breach reporting commitment, flow-down, return or destruction at termination, and the model training commitment, verified 8 August 2026
  • SiteGPT subprocessor list for every subprocessor and its stated location, and the zero-data-retention endpoints, verified 8 August 2026
  • SiteGPT security page for the SOC 2 Type II result, encryption, role-based permissions, and the trust portal, verified 8 August 2026
  • SiteGPT HIPAA workspace documentation for the seven day redaction default, lead handling, disabled sources and integrations, and escalation behavior, verified 8 August 2026

Last updated: August 2026. HHS guidance and all SiteGPT program pages were read directly on 8 August 2026.