The jagged frontier: when AI excels at maths and fails to read a clock
In short.
The best models win gold at the Mathematical Olympiad and misread an analogue clock one time in two. This is not a teething problem but a structural property. It explains why adoption reaches 88% while deployment in production on judgement tasks remains marginal.
The best AI models win gold at the Maths Olympiad, yet misread an analogue clock one time in two. The Stanford AI Index 2026 documents this paradox and its consequences for companies deploying AI on sensitive tasks.
A large language model can win a gold medal at the International Mathematical Olympiad. The same model, faced with an analogue clock, will tell the time correctly only one time in two.
This paradox has a name: the jagged frontier. It is now documented by the ninth annual AI Index report, published in April by the Stanford Institute for Human-Centered AI. And it changes what can be expected of AI in business.
A structural flaw, not a teething problem
Since 2017, the AI Index Report has been the neutral reference most cited by governments and companies on the real state of artificial intelligence. The 2026 edition, more than 400 pages across nine chapters, is signed by an independent committee bringing together academics, industry figures and public experts.
Its main finding fits in one sentence: large models are brilliant at times and failing at others, with no way of anticipating which.
The clock image is no accident. It illustrates what researchers call the jagged frontier: the same model excels at tasks humans consider extremely difficult and stumbles on tasks they consider trivial. No apparent logic, no reproducible pattern. The vendor itself does not know where the next error will occur.
For an executive, the consequence is less technical than strategic. A consumer tool cannot be deployed as is on sensitive tasks (compliance, contracts, audit, legal) because its failure zone is unpredictable. This is no longer an argument for caution: it is a measured fact, backed by figures.
88% adoption, fewer than one company in ten in production on judgement tasks
The report’s second figure challenges the prevailing narrative. According to Stanford, 88% of companies say they have “adopted” AI. But fewer than one in ten has actually put it into production on judgement tasks.
The deployments that work are concentrated on repetitive tasks (customer support, code generation, marketing), with productivity gains of 14 to 26%. On tasks that call for judgement, the measured effects are small, or even negative.
Yet compliance, legal monitoring, internal control and responding to audits are judgement activities by nature. The gap between the prevailing story of the AI revolution and its operational reality is now documented.
The silent collapse of transparency
The third finding directly concerns legal departments and audit committees. The foundation model transparency index calculated by Stanford fell from 58 to 40 points in one year.
Vendors are publishing less and less about what goes into training their models, the volume of data used, or the nature of the safeguards applied. For a data protection officer or an audit committee, the equation is simple: you cannot audit what you cannot see. And in the event of an incident, final liability remains with the user company, not with the vendor.
For the first time, the 2026 report also devotes an entire chapter to AI Sovereignty. A striking finding: public trust in US AI regulation has fallen to 31%, the lowest score among the major economies measured. Conversely, the European Union is now seen as the most credible region, a strong signal at a time when NIS2, DORA and the AI Act are reshaping the compliance obligations of large enterprises.
“A score of 75% on a legal reasoning benchmark says nothing about performance in a real law firm.” Raymond Perrault, co-director of the Stanford AI Index, April 2026
The statement deserves to be taken seriously. It refocuses the debate on three subjects that are no longer the business of the IT department alone: operational reliability, auditability and sovereignty.
The opposite of the jagged frontier
Optivalue.ai makes exactly the opposite bet: a specialised AI solution for the tasks where general-purpose models fall short (security questionnaires, compliance, ESG and tenders).
On reliability
Where consumer models show unpredictable reliability, Optivalue.ai relies on 85 agents trained vertically on domain-specific corpora. It is specialisation, not universality, that reduces the error surface.
On abstention
Where a general-purpose model always answers, even when it does not know, Optivalue.ai abstains when it has no source. It is the only AI that knows how to say “I don’t know”. Every answer cites its document, page and date. Zero hallucinations, by design.
On sovereignty
Where the transparency index of large models is collapsing, Optivalue.ai offers a private AI for each client, never pooled, available in a European sovereign cloud or on-premise. Your data never leaves your perimeter. ISO 27001 certified, winner of the European Sovereignty Award 2026.
The next analyses in this series will return to two practical consequences of the Stanford findings: the appearance of hallucinations on the most sensitive judgement tasks, and the confidentiality of data entrusted to large models, which has become a governance issue.
This analysis is based on a reading of the Stanford AI Index Report 2026, published in April 2026 by the Stanford Institute for Human-Centered AI. The report and its data are freely available at hai.stanford.edu/ai-index/2026-ai-index-report.
To find out how Optivalue.ai answers security questionnaires, audits and tenders without hallucinations: optivalue.ai
The editorial team, Optivalue.ai
Back to topOn the same topic