A New Chapter: Compliance & Risks is now Adherent. Read more.

7 min read
Blog

Why “Just Ask an LLM” Falls Short for Regulatory Coverage

Adherent (formerly Compliance & Risks). Abstract digital interface with glowing, colorful lines representing data flow, surrounded by various charts, percentages, and binary code on a dark blue grid background.

This blog was originally posted on 16th July, 2026. Further regulatory developments may have occurred after publication. To keep up-to-date with the latest compliance news, sign up to our newsletter.

AUTHORED BY STEPHEN BUGGY, ENGINEERING MANAGER, ADHERENT


Ask a general-purpose LLM which regulations apply to your product and you will get a confident, well-written answer in seconds. It is a tempting shortcut. It is also, for compliance purposes, a risky one. A confident answer is not the same as a correct one, and in regulatory compliance the cost of a missed rule is measured in recalls, fines and blocked shipments.

We wanted to know exactly how wide the gap is between a general-purpose LLM and a purpose-built agentic AI compliance platform. So rather than argue the point, we tested it.

Table of Contents

What We Set Out to Measure

The quality of any regulatory coverage tool comes down to two separate questions:

  1. Does it find everything that applies to my products?
  2. Is what it finds actually correct and up to date?

These are different problems and we measure them independently.

The first is a coverage question. If a platform surfaces ten regulations when forty apply, the thirty it missed are the ones that create risk because you never knew to look for them. The second is an accuracy question. A tool that identifies a hundred regulations is useless if half of them do not actually apply to your product, or point to law that has since been repealed or replaced.

A general-purpose LLM has no reliable way to answer either question about its own output. It cannot tell you what it missed, and it cannot independently verify what it returned is complete, current, and correct. That is the heart of the problem, and it is what our testing set out to measure.

How We Kept the Test Fair

To measure coverage we ran a like-for-like benchmark. We took a product that had already been assessed in Adherent, then put the same product and the same market to three leading LLMs : ChatGPT, Claude and Gemini, each queried off the shelf.

Fairness mattered, so we held the conditions steady:

  • The same input a customer would give. Each model received only a plain product description, for example “a battery-powered chainsaw sold to consumers in Germany.” Nothing more.
  • No head start. We did not hand the models any of the structured profiling that Adherent builds internally. Giving that away would have flattered the result and told us nothing useful.
  • One question, once. A single query per model, no coaching, no retries to fish for a better answer.
  • A like-for-like count. Every regulation a model named was matched against Adherent’s list by its official identifier, such as an EU CELEX number or a national statutory instrument number.

We then repeated the exercise across several product types and several markets, from the EU as a whole, down to individual member states , and, separately in the United States.

What the Benchmark Found

The pattern held in every market we tested. Adherent identified many times more applicable regulations than any single public model, and the models repeatedly surfaced the same well-known regulations while missing the majority of applicable requirements.

ProductMarketAdherent foundGeneral-purpose LLMs 
Bluetooth speakerEU3527 to 20
Bluetooth speakerIreland6515 to 26
Battery-powered chainsawGermany8912 to 19
Battery-powered drillUnited States1148 to 13
Wi-Fi robot vacuumUnited States1578 to 23

The EU run is the starkest example. Adherent identified 352 applicable regulations for a single Bluetooth speaker. The best-performing public model found 20. Approximately  300 applicable rules were missed by every model we tested. All three agreed only on the five best-known EU regulations, covering radio equipment, hazardous substances, waste, chemicals and batteries.

In the United States, the overall regulatory framework is similar across products, but much of the critical missing detail sits at the state level.  For a Wi-Fi robot vacuum, Adherent tracked 157 applicable regulations across federal and state law. The general-purpose LLMs returned between 8 and 23, converging on the familiar federal safety canon and one or two headline California laws while missing the fast-moving state-level rules on privacy, product stewardship and automated decision-making.

Why the Gap Is So Large

The models are not necessarily wrong about the laws they identify. The problem is that they often stop at the most visible layer of regulation. That is now where compliance is determined. 

In the EU, a directive rarely creates the requirements a manufacturer must follow directly. Those requirements are established through the national laws that implement the directive, together with later amendments, technical rules, and other implementing legislation that define what compliance requires in practice.

Those deeper layers account for much of the gap between the public model results and Adherent’s findings. Two findings make the risk clear:

  • The models cite laws that are no longer in force.
    In the EU analysis, two of the three models identified a product safety directive that had already been repealed and replaced. Adherent identified the current regulation instead.
  • The models are not consistently able to produce an answer.
    In the Ireland analysis, one model failed to return a usable response across five separate attempts, despite answering the same prompt at the EU level. An answer you cannot rely on is not a compliance tool.

Single-country analysis exposes another challenge. The models often identify the parent EU directive, while the requirements that actually apply are defined in national law, such as an Irish statutory instrument introduced in 2024. Because the directive and the national law rarely share the same identifier, a compliance professional is still left to find and verify the connection the model missed.

The issue is not whether a model can name a relevant law. It is whether it can identify the current law that establishes the requirements for a product in the market where it is sold.

Coverage Is Only Half the Story

Finding more regulations is only valuable if the applicability decisions behind them are sound. This is where our quality approach differs most from simply trusting a general-purpose LLM .

Adherent uses purpose-built AI agents to assess whether a regulation applies to a specific product profile. Every decision is validated against an expert-curated reference dataset built over nearly 25 years by regulatory subject matter experts. Rather than relying on a model’s own confidence, the agents are evaluated against thousands of real applicability decisions made by compliance experts. That validation is continuous and built into the platform itself, not something performed only for benchmarking.

General-purpose LLMs have no comparable validation framework, and by their nature it is difficult to create one. They can generate plausible answers, but they are not evaluated against decades of expert applicability decisions, nor can they consistently demonstrate how those decisions were reached. For a regulatory register that must withstand scrutiny from an auditor, regulator, or customer, traceability matters as much as the answer itself. If you cannot trace why a requirement was identified and validate the reasoning behind it, you cannot rely on it. 

What This Means for Your Team

The benchmark makes the risk tangible. For a single product in a single market, the difference between a general-purpose AI model and a platform purpose-built for product compliance was hundreds of missed requirements, including current laws replaced by legislation that is no longer in force. In compliance, the requirement you never identified is often the one that becomes an incident.

That is why we treat coverage and accuracy as first-class metrics, continuously validating our AI agents against expert (human) judgement. 

If you’d like to see how the coverage for your own products compares with the tools you rely on today, we’d be happy to walk you through the results.

See Adherent in Action

Discover how agentic AI is reshaping product compliance for global enterprises.

Sign up to our

Monthly Market Insights

Our Connect Newsletter delivers the latest regulatory developments, trends and expert insights straight to your inbox.