← Glossary

Golden set (AI agents)

A set of procedures and data that are designed, repeatable and prepared to compare a system’s current result with a reference result considered correct. The AEPD guidelines on agentic AI include it as a “golden testing” practice within the continuous, evidence-based evaluation of the agent (section VII.B, pp. 57-58); the reference result is called a golden result or golden sample.

It is the test that is repeated with the same cases to see whether the agent still responds as it should after a change.

Which obligations it carries

It is a best practice, not a legal obligation: it is one of the validation testing techniques proposed by the guidelines. It serves obligations that are legal: that personal data be accurate and, where necessary, kept up to date (Article 5(1)(d) of Regulation (EU) 2016/679) and that the controller be able to demonstrate that it respects the principles (Article 5(2)), both applicable from 25 May 2018.

What it is not

It is not a one-off test before deployment: its value lies in being repeated, because it allows repeatable checking and the assessment of deviations in the face of changes in legal terms and in the functionality of systems (p. 58). Nor is it an A/B test or a benchmark, which the guidelines mention as different evaluation methods (p. 57). And it is not enough on its own to repeat a specific decision: for that the guidelines also propose repeatability mechanisms, such as keeping a record of the configuration in a decision process (section VII.F, p. 69).

The nuance almost nobody captures

A golden set is only as good as its reference cases: if it does not include the situations that affect people — a request to exercise rights, an inaccurate data item, a special category of data — it will confirm that the agent is unchanged where it matters least. And since the guidelines warn that, without strict control over sources, services, their versions and memory, the output of an agentic system cannot be anticipated (section IV.D, p. 30), the comparison with the reference result has to set what difference is tolerated.

To find out more

Reviewed on 11 October 2026. Dates according to Article 113 of Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744 (OJ of 24 July 2026, in force since 27 July 2026).