ΑΙhub.org
 

AAAI presidential panel – AI evaluation


by
21 September 2026



share this:

Three green rabbits sit at the bottom of the image with dotted lines in a square around them, with an "i" information symbol captured within a speech mark. In the background is a sketch, set against a purple background. A neural network diagram is overlaid the background. Marcin Wilkowski / More info please / Licenced by CC-BY 4.0

The Future of AI Research report, published in March 2025, aims to clearly identify the trajectory of AI research in a structured way. The report was led by outgoing AAAI President Francesca Rossi and covers 17 different AI topics. Members of the report team, and other selected AI practitioners, are taking part in a series of video panel discussions covering selected chapters from the report.

In the next discussion in the collection, the panellists discuss AI evaluation. Specifically, they cover the following topics:

  • Why standard software testing methods fail for AI systems
  • The limits of benchmark-driven testing (and how Goodhart’s Law creeps in)
  • The four dimensions of evaluation the report says we’re missing: capability, usability, human-system performance, and legal/ethical compliance
  • Open research challenges — from monitoring systems that evolve after deployment to evaluating agentic AI safety

Panel Members

  • Dawn Song, University of California, Berkeley
  • Sara Hooker, Co-Founder, adaption
  • Karen Myers, Carnegie Mellon University

Moderator

  • Francesca Rossi, AAAI past president, IBM Fellow and AI Ethics Global Leader


tags: ,


Lucy Smith is Senior Managing Editor for AIhub.
Lucy Smith is Senior Managing Editor for AIhub.

            AUAI is supported by:



Subscribe to AIhub newsletter on substack



Related posts :

Disappearing lakes and AI are helping scientists map Arctic permafrost thaw in near‑real time

  18 Sep 2026
Researchers created an interactive website to track permafrost thaw across the Arctic.

How much can fair budget-division rules resist manipulation?

The authors write about their award-winning IJCAI-ECAI paper: "Approximate Strategyproofness in Approval-based Budget Division".

AI dives into a sea of data, from plankton to pollution

  16 Sep 2026
“Faster and cheaper monitoring means problems like plankton decline, litter accumulation, oil spills and coral degradation can be picked up and acted on sooner."

Interview with Yash Saxena: how is external knowledge used in AI systems?

  15 Sep 2026
What happens to information from external sources as it moves through an AI system?

When AI art has no author: Study finds generated images often can’t be traced to training data

  14 Sep 2026
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.

AI in nature conservation: powerful tool or dangerous shortcut?

  11 Sep 2026
AI provides opportunity for future biodiversity conservation but introduces risks .

Improving the process for large-scale recommender systems: an interview with Haruka Kiyohara

  10 Sep 2026
Credit-assigned policy gradient for early stage retrieval in two-stage ranking.

The Machine Ethics podcast: Data Collective with E.M. Lewis-Jong

Ben chats to E.M. Lewis-Jong about the promise of AI and making human connection easier, speech recognition and supporting linguistic diversity, making useful technologies that have a purpose, and more.



AUAI is supported by:







Subscribe to AIhub newsletter on substack




 















©2026.05 - Association for the Understanding of Artificial Intelligence