AI Summaries Score 47% in Government Trial, Humans Outperform at 81%
by Joy Veyra 2026-08-23

AI Summaries Score 47% in Government Trial, Humans Outperform at 81%

Compiled by the editorial desk with reference to the trial report and public statements from ASIC and Amazon Web Services.

In a head-to-head test conducted for an Australian financial regulator, generative AI produced document summaries that scored a dismal 47% on accuracy, while human employees scored 81%. The trial, run by Amazon Web Services for the Australian Securities and Investment Commission (ASIC), used Meta's open-source Llama2-70B model to summarize real government documents submitted to a parliamentary inquiry.

The results, first reported by Australian outlet Crikey, underscore a growing concern among organizations experimenting with generative AI: the technology often fails to deliver reliable, useful output in professional contexts. ASIC commissioned the trial as a proof of concept to assess whether AI could handle business-related summarization tasks, but the findings suggest the technology is not yet ready for such roles.

How the Test Was Conducted

Human employees at ASIC and the AI model were given the same set of documents and asked to produce summaries that focused on ASIC-related content, including references and page numbers. Five evaluators then read the original documents and assessed the summaries blindly—they were labeled only as A and B, with no indication that one set was generated by AI.

After the evaluations were complete, three of the five assessors said they suspected they had been reviewing AI outputs. The AI performed worse than the human summaries across all criteria, according to the report.

One notable failure was the AI's inability to provide accurate page numbers for its citations—a flaw that the report suggests could be corrected with adjustments to the model. But more fundamental issues emerged: the AI frequently missed nuance and context, made puzzling choices about what to emphasize, and included irrelevant or redundant information. The summaries were described as "waffly" and "wordy."

Implications for Workplace Use

The assessors concluded that using AI summaries would likely require significant fact-checking, which could negate any potential time or cost savings. This finding challenges the narrative that generative AI can streamline business operations without sacrificing quality.

The trial's results echo broader critiques of generative AI's reliability, particularly in tasks that require careful judgment and attention to detail. While the Llama2-70B model is not the newest available, its 70-billion-parameter architecture is considered capable, making the poor performance all the more striking.

For organizations considering AI adoption, the ASIC trial offers a cautionary tale. The technology may excel in narrow, well-defined tasks, but for complex document analysis, human expertise remains superior. The report's findings could influence how regulators and businesses approach AI integration in the near term.

AI Summaries Score 47% in Government Trial, Humans Outperform at 81%

A government-commissioned trial in Australia found that AI-generated summaries scored only 47% on accuracy, while human summaries scored 81%. The AI, based on Meta's Llama2-70B, struggled with page references and nuance, raising doubts about its practical use in business settings.

Leave a Comment

Comments (0)