RatedWithAI

RatedWithAI

Accessibility scanner

AI LiabilityAugust 19, 2026

AI Product Safety 2026: When a Model Update Becomes a Reportable Defect

Consumer product safety law obliges manufacturers to report defects that could create a substantial hazard, promptly, before anyone establishes that the product is unsafe. On AI-enabled products the design ships weekly, changes after sale, and is triaged by engineers who have never read that standard.

Report ≠ admit
The duty triggers on information, not on a conclusion that the product is defective
OTA ≠ done
A pushed fix does not reach the unpatched tail or undo the injuries
Perishable
The model version that caused it may not exist by the time anyone asks

The Product Now Changes After You Sell It

Product safety regulation assumes a stable artefact. A unit is designed, tested, certified, manufactured and sold, and the thing in the customer's home is the thing that left the factory. Everything downstream — the certification, the test report, the instruction manual, the hazard analysis — describes that artefact.

AI-enabled consumer products break the assumption at its root. A robot vacuum's obstacle behaviour, a smart oven's cook logic, a monitor's alerting thresholds, a fitness device's guidance, a connected toy's conversational scope: all of these are defined by models that can be replaced remotely, sometimes weekly, sometimes on a rolling percentage of the fleet. The unit in the customer's home is no longer the unit that was tested, and the company frequently cannot say which version any given unit is running without querying a deployment system.

This produces a specific organisational failure that has nothing to do with anyone behaving badly. Model regressions are handled by the machine-learning team as quality issues, tracked in the same tooling as latency and accuracy, and closed when the metric recovers. Nobody in that loop is asking the regulatory question — does this fault create a risk of injury such that reporting is required — because the regulatory question lives with the product-safety function, and no ticket routes there.

A Defect Taxonomy for Model-Driven Products

Four categories, each of which maps onto an established defect theory. The point of the taxonomy is to give safety-relevant model failures a name that a product-safety reviewer recognises, so the ticket routes correctly.

Perception failure

Design defect

The model fails to recognise a hazard state it must recognise: a person in the path, a child's hand, a spill, an obstruction, a temperature excursion. Frequently correlated with conditions that are unevenly distributed across users — lighting, skin tone, floor colour, accent, ambient noise — which turns an accuracy gap into an unevenly distributed injury risk.

Actuation failure

Design defect

The model correctly perceives and then acts unsafely: continuing to heat, failing to yield, applying force, resuming after an interruption it should treat as a stop condition. This is the category where the physical harm pathway is shortest and where a hardware interlock, independent of the model, is usually the only adequate control.

Advice failure

Warnings and instructions defect

An assistant emits safety-relevant guidance that is wrong or contradicts the manual: a cooking temperature, a cleaning-agent combination, a maintenance step performed with power connected, an age-appropriateness judgement. Warnings and instructions are an express defect category, so this is not a peripheral exposure just because no actuator moved.

Post-sale behaviour change

Design change without re-evaluation

An update alters behaviour the original hazard analysis relied on — a threshold moved, a conservative default relaxed, a feature extended to a new context. The certification and test report still describe the previous behaviour, and no one re-ran the hazard analysis because the change was framed as a model improvement rather than a design change.

The Clock, From Signal to Filing

The reporting standard is deliberately low and deliberately fast: information reasonably supporting the conclusion that a defect could create a substantial hazard, evaluated promptly, reported without waiting for certainty. What follows is the internal path that has to exist for that standard to be met, expressed as elapsed time from the first signal.

1

Day 0 — Signal

A support ticket, app review, telemetry anomaly, warranty claim, retailer complaint or internal test result touches a safety-relevant behaviour. In practice the signal almost never arrives labelled as a safety issue.

2

Day 0–2 — Routing

The signal reaches the product-safety function, not only engineering. This requires a keyword and category screen across every intake channel, because the routing failure — not the analysis — is what makes reports late.

3

Day 2–5 — Reproduction and scoping

Identify the model version and configuration involved, reproduce the behaviour, and query the deployment system for how many units are running affected versions and where. Scope is what turns an anecdote into a hazard assessment.

4

Day 5–10 — Hazard evaluation

Apply the regulatory standard rather than a bug-severity rubric: severity of possible injury, likelihood, exposed population, and whether the user can perceive and avoid the hazard. Record who performed it and on what facts.

5

Within the statutory window — File

Report if the standard is met, including where investigation is ongoing. A report filed while investigating is expressly contemplated; a late report from a company that was still deliberating is the enforcement pattern.

6

Continuing — Remedy and follow-through

Corrective action covers the unpatched tail, not just the fix: forced-update mechanics, feature disablement, direct notice to registered owners, retailer notification, and a plan for units that never reconnect.

What Has to Be Built Before the First Incident

  • Version-to-unit traceability. You must be able to state which model version any unit was running on a given date. Without it you cannot scope a hazard, cannot design a remedy, and cannot answer the first question anyone will ask after an injury.
  • A safety screen on every intake channel. Support, app-store reviews, warranty, social, retailer complaints and telemetry all need a route to the safety function. Most companies have one channel wired and assume the others feed it.
  • Model changes treated as design changes. Any update touching a behaviour the hazard analysis relied on gets a documented re-evaluation before rollout — the same discipline applied to a mechanical change.
  • Extended log retention for safety-relevant events. Default telemetry retention is set for engineering convenience and is far shorter than the window in which a claim arrives. Safety-relevant events need their own schedule.
  • Model-independent interlocks. Where an unsafe actuation is possible, the control that prevents it should not be the same model that might fail. This is the difference between a recall and an incident report.
  • A constrained answer path for safety questions. If your assistant can be asked about temperatures, chemicals, dosing, children or maintenance, those domains need retrieved, approved answers rather than open generation.

Frequently Asked Questions

We license the model from a vendor. Does the reporting duty sit with them?

The duty sits with the parties the statute names — typically manufacturers, importers, distributors and retailers of the consumer product — and a model vendor supplying a component is generally not the entity placing the product on the market. You are. That has two consequences worth acting on now. Contractually, your agreement should require prompt notification of safety-relevant defects, regressions and evaluation failures, on a clock that lets you meet your own; a vendor learning about a perception failure and telling you in the next quarterly review has made you late. Practically, you need the ability to pin, roll back and reproduce a specific version, which is a technical requirement to negotiate before signing rather than a favour to request during an incident. A vendor who cannot tell you what changed between versions cannot support a product-safety obligation you cannot delegate.

Is a report an admission that our product is defective?

No, and the frameworks are built to make that explicit because they need companies to report early. The standard is information reasonably supporting a conclusion, which is well below proof, and reports are routinely filed and closed without a recall. Firms are generally permitted to report while stating that they do not concede a defect exists, and the reporting form contemplates ongoing investigation. The asymmetry is what should drive the decision: reporting something that turns out to be nothing costs an investigation and some staff time, while failing to report something that turns out to be real produces civil penalties assessed on the delay itself, on top of whatever the underlying issue costs. Late reporting is the enforcement action companies most reliably suffer, and it is entirely self-inflicted.

How should we handle a defect that only appears for some users?

Treat unevenly distributed failure as an aggravating factor rather than a limiting one, because that is how it will be read. A perception model that performs worse in low light, on dark flooring, with certain skin tones, with accented speech or in smaller-statured users produces a hazard concentrated on an identifiable population. For the safety analysis, the relevant figure is the risk to the exposed subgroup, not the fleet-wide average — a failure rate that looks acceptable across all users can be unacceptable for the people who experience it. Separately, uneven safety performance has a second life outside product-safety law entirely, in consumer-protection and civil-rights frameworks that ask whether a product performs materially worse for a protected group. Document subgroup performance during evaluation, because the absence of that data is itself a finding once someone goes looking.

Our product is sold in multiple countries. Does one report cover it?

No. Product safety reporting is jurisdictional, and the triggers, clocks, thresholds and recipients differ — as do the recall mechanics and the market-surveillance authorities. A defect in a model deployed to a global fleet is by construction a multi-jurisdiction event on day one, and the practical consequence is that your incident procedure needs a jurisdiction matrix prepared in advance rather than assembled during a live issue. Two further points that catch connected-product companies: some regimes now impose their own reporting duties for serious incidents involving AI systems, which can run in parallel with the product-safety duty on a different clock and to a different authority; and geographic deployment rings mean the affected population may not match your sales geography at all, so scope the report to where the version actually shipped.

Can we avoid all this by disabling the AI feature remotely?

Remote disablement is a legitimate and often excellent corrective action, and it is not a substitute for the report. It also introduces its own set of problems that should be worked through before you rely on it. Disabling a feature customers paid for is a consumer-protection and contract question, particularly where the feature was central to the purchase decision, and it may create an obligation to refund or compensate. It only reaches connected units, so the offline tail persists exactly as it does for a fix. And if the feature is load-bearing for the product's basic function, disabling it may create a different hazard — a monitoring product that stops alerting is not obviously safer than one that alerts imperfectly. Plan disablement as one branch of a documented remedy, alongside notice to owners, a route for unconnected units, and the filing itself.

The Question That Reveals the Gap

Pick one shipped unit and ask your team which model version it is running today, which it was running ninety days ago, and who evaluated the safety impact of the change between them.

If the first two answers require a week of work and the third has no owner, the reporting duty cannot be met on time — not because anyone would refuse, but because the facts needed to meet it do not exist yet.