Journalists need their own benchmark tests for AI tools

The performance tests used by AI companies don’t measure what matters in the newsroom

Posted

A recent paper from OpenAI researchers sheds new light on why large language models (LLMs) are prone to “hallucination,” or fabricating information. According to the paper, the evaluation methods major AI companies use encourage overconfidence. Performance tests often take the form of multiple-choice questions with explicit correct answers that end up unintentionally rewarding models for guessing rather than declining to answer if they aren’t certain. By optimizing their systems to achieve a high score on these evaluations, AI companies are training their models to be good test-takers instead of actually improving their overall accuracy.  ...

Most AI tools aren’t designed with journalists or news audiences in mind, and benchmarks used by AI companies rarely measure what matters in the newsroom. As a result, reporters, editors, and fact-checkers lack visibility into whether the ever-evolving models are suited to their needs, or how their outputs stack up against journalistic values like accuracy, transparency, accountability and objectivity. As Charlotte Li, a computational-journalism PhD student at Northwestern University, puts it, “Do any of these scores tell us which models we should use for journalism and when?”

Recognizing this knowledge gap, the Generative AI in the Newsroom project, led by Nicholas Diakopoulos, the director of the Computational Journalism Lab at Northwestern University, is pushing for the development of benchmarks tailored to journalism.

Read more from Columbia Journalism Review