LONDON – For three years, the Bank of England set interest rates partly on employment data it quietly suspected was unreliable. Survey response rates had collapsed during the pandemic, the numbers arriving at Threadneedle Street carried wider error margins than policymakers acknowledged publicly, and when the Office for National Statistics disclosed last year that an operational failure had caused interviewers to be under-allocated across key geographic areas, it confirmed what economists had long whispered: Britain’s flagship labour market measure had a credibility problem.
James Benford, Director General for Surveys and Economic Statistics at ONS, said this week that artificial intelligence now sits at the centre of the recovery effort. Bloomberg reported on Monday that the agency believes AI can help it repair statistical integrity that took years to erode – a judgment that marks a significant shift in how a traditionally conservative statistical office describes its own modernisation.
“Putting quality ahead of quantity, and credibility ahead of volume, remains a guiding principle,” Benford wrote in an ONS July 15 progress update, the most candid public accounting of the recovery effort since data problems began surfacing in earnest. The document described operational improvements, expanded staffing, and technology deployments that the director general said have placed the agency on what he called a “firmer footing.”
The centrepiece tool is ClassifAI, an in-house system that automates the classification of occupations from survey responses. The manual process – cross-referencing job descriptions against the Standard Occupational Classification – previously consumed significant analyst hours. ClassifAI has halved occupation-coding processing time, according to ONS, releasing staff to focus on validation and quality assurance tasks that automation cannot yet perform. The efficiency gains feed directly into the employment breakdowns that underpin UK inflation data models and Bank of England deliberations.
The Labour Force Survey itself – the monthly snapshot that feeds every major UK employment indicator – has seen response rates recover to broadly pre-pandemic levels. But the agency frames this as stabilisation rather than vindication. Its replacement, the Transformed Labour Force Survey, is generating data showing improved responses – particularly in later survey waves – and remains on track to become the headline employment measure by 2027. The transition requires simultaneous maintenance of two competing instruments across the same population, a logistical exercise the ONS has not previously attempted at scale.
The range of AI applications extends beyond occupation-coding. ONS has integrated grocery scanner data into the construction of its Consumer Prices Index, replacing slower and more error-prone manual price collection for a significant segment of the household basket. Automated receipt-reading tools, part of a broader Reproducible Analytical Pipeline, are now processing family spending diaries from a household sample expanded to more than 30,000 participants. ONS says the grocery integration has roughly halved the error rate compared to the previous methodology. The agency also added 214 statistical staff during the period under review, headcount expansion that sits alongside the automation push rather than in place of it.

The ONS’s AI push has a certain recursive quality. According to its own business surveys – the same type of data collection whose reliability is under repair – 29% of British companies now use artificial intelligence in their operations, rising to almost 50% among the largest firms. That figure is itself an object lesson in the stakes of what the agency is trying to fix: if the survey underpinning a key technology-adoption metric has reliability problems, the statistic it produces carries those problems forward. The broader AI hardware demand surge that has reshaped global trade over the past year makes accurate technology-adoption statistics economically significant well beyond Britain.
Benford’s July progress update included a disclosure that went further than standard agency communications. He named the interviewer under-allocation failure directly, releasing a corrected tally: 17 lower-impact statistical errors identified and corrected between March and May 2026, with no corrections to higher-impact errors during that window. “All tier 1 economic statistics were published on time,” the blog post stated. The disclosure was calibrated as a transparency gesture, but it also introduced a level of error classification granularity the agency had not previously applied to public communications.
The International Monetary Fund has assessed ONS at its highest standards tier for statistical accreditation purposes. The evaluation applied IMF Article IV and SDDS Plus criteria; whether it encompasses the LFS disruption period or applies primarily to the more stable core national accounts and price indices has not been specified in IMF documentation. Several economists with knowledge of the process note that the accreditation decision was taken before the interviewer allocation failure came fully to light – without suggesting bad faith on either side.
The ONS effort is taking shape inside a policy environment in which artificial intelligence fraud and governance questions have moved up every government agenda simultaneously. The government AI governance frameworks under construction in Washington have created a policy context in which statistical offices that do not automate risk falling behind the economies they are supposed to measure. For the ONS, which spent the better part of three years explaining why its most-watched indicator had become less trustworthy, the AI pivot also carries an institutional reputational logic: technological modernisation is easier to communicate publicly than the slower work of rebuilding survey infrastructure.
The uncertainty in Benford’s account concerns less what has been fixed than what remains untested. ClassifAI addresses occupation misclassification – a processing error – but not the non-response bias that made LFS data unreliable during and after the pandemic. Those are different failure modes. When low-response surveys return predominantly from households that happen to answer, the resulting sample is no longer representative regardless of how accurately the recorded occupations are classified. The TLFS, once it becomes the headline measure, will also introduce a break in the time series. ONS has not yet released the splicing methodology that analysts will need to maintain backward comparability across the transition.
A fuller AI development plan – covering the next wave of automation beyond ClassifAI and the receipt-reading pipeline – is to be finalised by the end of 2026. What it will target, how it will be audited, and what decision-making processes will govern AI-generated outputs in official statistics have not been disclosed. AI cooperation between governments has yet to produce binding standards for official statistics specifically, leaving the Office for Statistics Regulation – which holds formal oversight authority over ONS – to develop its own review framework without a settled international template. That oversight assessment, when it comes, may be the most consequential test of whether the ONS technology turn amounts to a genuine statistical repair or a well-timed rebranding.

