ΑΙhub.org
 

New algorithm follows human intuition to make visual captioning more grounded

visual captions | AIhub

Annotating and labeling datasets for machine learning problems is an expensive and time-consuming process for computer vision and natural language scientists. However, a new deep learning approach is being used to decode, localize, and reconstruct image and video captions in seconds, making the machine-generated captions more reliable and trustworthy.

To solve this problem, researchers at the Machine Learning Center at Georgia Tech (ML@GT) and Facebook have created the first cyclical algorithm that can be applied to visual captioning models. The model is able to use the three-step processing during training to make the model more visually-grounded without human annotations or introducing additional computations when deployed, saving researchers time and money on their datasets.

The algorithm employs attention mechanisms, an intuitive concept for humans, when looking at a photo or video. This means that it tries to determine what aspects are important in an image and sequentially create a sentence explaining the visual.

This new model helps solve issues with previous attempts where an algorithm would make its decision based on prior linguistic biases instead of what it is actually “seeing.” This would lead to algorithms having what researchers refer to as object hallucinations. Object hallucinations occur when an algorithmic model assumes an object like a table is in a photo because in previous images, someone with a laptop was always sitting at a table. In this instance, the model is unable to understand a situation where a person has a laptop on their lap instead of a table. This new model helps alleviate the object hallucination problem, thus making the model more reliable and trustworthy.

Chih-Yao MaChih-Yao Ma, a Ph.D. student in the School of Electrical and Computer Engineering, envisions this model being used in situations like describing what happens in the scene as a technology to assist people who are visually impaired to overcome their real daily visual challenges. The model would be a good fit in such instances, because it can alleviate the linguistic bias and object hallucination issues in existing visual captioning models.

This work has been accepted to the European Conference on Computer Vision (ECCV), which takes place virtually August 23-28, 2020.

For more information on ML@GT at ECCV, visit our conference website.

Read the paper in full

Learning to Generate Grounded Visual Captions without Localization Supervision
Chih-Yao Ma, Yannis Kalantidis, Ghassan AlRegib, Peter Vajda, Marcus Rohrbach, Zsolt Kira
Georgia Tech, NAVER LABS Europe, Facebook




Allie McFadden is the communications officer for the Machine Learning Center at Georgia Tech and the Constellations Center for Equity in Computing at Georgia Tech.
Allie McFadden is the communications officer for the Machine Learning Center at Georgia Tech and the Constellations Center for Equity in Computing at Georgia Tech.

Machine Learning Center at Georgia Tech

            AUAI is supported by:



Subscribe to AIhub newsletter on substack



Related posts :

The Machine Ethics podcast: Safe and moral AI with Rebecca Raper

In this episode Ben chats with Rebecca about AI governance and guardrails, moral assurance, under-specification problems, lack of interdisciplinary work in robotics, and more.

CogTwin: A framework for adaptable digital twins

  05 Aug 2026
Find out more about work presented at IJCAI 2025 on Cognitive Digital Twins.

Forthcoming machine learning and AI seminars: August 2026 edition

  04 Aug 2026
A list of free-to-attend AI-related seminars that are scheduled to take place in the next couple of months.

Healthcare benchmarks are only as good as their assumptions

  03 Aug 2026
In healthcare settings where patients use LLMs as a medical assistant, LLM performance differs between evaluation and deployment.

Engineering Out Loud: S13E2 – Ethics in AI presentation

  31 Jul 2026
Hear from Oregon State University researchers Houssam Abbas and Alicia Patterson.

Humans trained to spot AI faces in the battle against deepfake fraud

Researchers trained people to spot AI-generated faces by drawing their attention to six perceptual qualities.
monthly digest

AIhub monthly digest: July 2026 – time-series anomaly detection, music generation, and RoboCup in action

  29 Jul 2026
Welcome to our monthly digest, where you can catch up with AI research, events and news from the month past.

OpenAI’s models autonomously hacked a tech startup. It signals a seismic shift in cybersecurity

  28 Jul 2026
An autonomous agent powered by OpenAI’s went rogue during a security test and hacked multi-billion dollar tech startup, Hugging Face.



AUAI is supported by:







Subscribe to AIhub newsletter on substack




 















©2026.05 - Association for the Understanding of Artificial Intelligence