RatedWithAI

RatedWithAI

Accessibility scanner

AI Legal & ComplianceAugust 9, 2026

You Called It an A/B Test. The Common Rule Calls It Research.

The experiment ships on a Tuesday. Six months later someone wants to write it up, a reviewer asks for the ethics approval number, and a question nobody asked in advance becomes unanswerable in retrospect.

Why AI companies hit this and ordinary software companies mostly did not

AI teams publish. The field's hiring, credibility and recruiting run on papers, preprints and blog posts making claims about how people behave with these systems — and the moment an internal experiment is reframed as a contribution to what is known generally, it satisfies the first half of a definition that most product teams have never read. The second half is satisfied almost automatically, because the underlying data is conversation.

Both Gates, or Neither Applies

Gate 1 — Research?

A systematic investigation, including development, testing and evaluation, designed to develop or contribute to generalisable knowledge.

Satisfied by
  • +Intent to publish, present or post a preprint
  • +Hypothesis stated in advance and tested against a control
  • +Conclusions framed as being about people in general, not about your users
  • +A named collaborator whose output is an academic paper
Generally not
  • Optimising a flow to raise your own conversion rate
  • Internal quality evaluation with no external claim
  • Operational monitoring and incident analysis
Gate 2 — Human subjects?

A living individual about whom the investigator obtains information through intervention or interaction, or obtains identifiable private information.

Satisfied by
  • +Manipulating what a user sees in order to measure their response
  • +Surveys, interviews and moderated sessions
  • +Analysing conversation transcripts tied to accounts
  • +Re-identifiable behavioural traces, however indirect
Generally not
  • Benchmarks run against models with no human data
  • Genuinely de-identified aggregate statistics
  • Public documents about organisations rather than people

The useful consequence of a conjunctive test is that one honest answer can resolve it. An experiment whose findings genuinely stay inside the company clears gate one; a study that truly touches no identifiable human data clears gate two. Most disputes are about teams asserting both while planning to publish.

What the Team Says, What the Definition Asks

"It's just an A/B test, we run hundreds."

Is anyone going to write about what it showed? Volume is irrelevant to the definition; the intended contribution to generalisable knowledge is the whole of gate one.

"The data is anonymous — we removed the user IDs."

Could identity readily be ascertained from what remains? Free-text prompts routinely carry employer, location, health and family detail that no ID-stripping touches.

"Users agreed to research in the terms of service."

Were the required elements of informed consent present, and was refusal genuinely without penalty? Acceptance of terms as a condition of access answers neither question.

"Our university partner is handling the ethics side."

Is your company engaged in the research — obtaining data, interacting with participants, or receiving identifiable information? If so you have your own obligations, not a borrowed exemption.

"It's quality improvement, not research."

Where does the knowledge land? The distinction is defensible when findings stay internal and operational, and collapses the moment the same analysis becomes a public claim about human behaviour.

"It's minimal risk, so review isn't needed."

Who determined that? Minimal risk changes the level of review and may support an exemption, but that determination is made by a review body, not by the investigator who wants the answer.

Engagement Is How the Obligation Reaches You

Companies routinely assume a university collaborator absorbs the compliance work. Sometimes that is right and sometimes it is exactly backwards. The concept that controls the answer is engagement: an entity is generally engaged in research when its own employees or agents intervene or interact with participants for research purposes, or obtain identifiable private information for those purposes. A company that runs the experiment inside its own product, holds the raw transcripts and hands a partner an analysis is doing far more of the regulated activity than the partner is.

The practical version of this is a written determination made before data collection — who is engaged, whose review covers what, and whether a reliance arrangement is needed — rather than an assumption resolved by whoever wrote the acknowledgements section.

Six Designs That Deserve Review Regardless

Even where the formal regulation does not bind you, a serious deployment review is worth running on the categories below — because the reputational and litigation risk does not track the funding source, and because the first external question after an incident is invariably who reviewed this before it ran.

01Studies involving users who are minors, or products where age is unverified and minors are foreseeably present.
02Anything touching mental health, crisis behaviour, self-harm signals or emotional manipulation — the highest-risk category and the one most attractive to an AI research agenda.
03Deception or incomplete disclosure designs, which require a specific justification and usually a debriefing plan.
04Experiments on employees, contractors or annotators, where voluntary refusal is structurally difficult.
05Health-adjacent inference from behavioural data, which can pull in a separate regulatory regime alongside this one.
06Any design where the control arm is meaningfully worse for a user with a real problem in front of them.

The Cheapest Version of Getting This Right

A one-page determination form attached to any experiment where publication is plausible, answering both gates and naming who decided, costs a team almost nothing and resolves the entire question at the only moment it is cheap to resolve. Independent review boards will review commercial protocols for a fee where a formal record is needed. Both options are dramatically less expensive than the alternative, which is discovering at submission that a year of work is unpublishable and that the dataset behind it was assembled under a consent theory nobody would defend in writing.

Related Reading

Check What Your Site Says About User Data and Research

"We never use your conversations", "opt out anytime" and research-programme pages accumulate across docs, trust centres and old launch posts — and they are the record against which a consent theory gets judged.

See every claim your site makes in one pass. Run a free scan and check each against what your experiments actually do.

This article is general information and not legal or regulatory advice. The scope of federal human-subjects regulation depends on funding, institutional assurances and the specific facts of a study, and institutional policies frequently impose more than the regulation requires. Consult qualified counsel or a review board before relying on any conclusion here.