{"692819":{"#nid":"692819","#data":{"type":"news","title":"New Benchmark Measures Memory Capabilities of VLM-Powered Robots","body":[{"value":"\u003Cp\u003EIn-home assistive robots might be able to clean and organize your home, but if they can\u2019t tell you where they put your keys or wallet, they can become a major inconvenience.\u003C\/p\u003E\u003Cp\u003ETo improve robot performance and advance memory capabilities needed for reliable home use, Georgia Tech researchers have developed a new benchmark to evaluate the memory capabilities of vision-language models (VLMs).\u003C\/p\u003E\u003Cp\u003EThat includes Gemini 2.0-flash and GPT-4o.\u003C\/p\u003E\u003Cp\u003EIn-home assistive robots could clean and organize a home while the owner is away, but without enhanced memory capabilities, they could cause more problems than they solve.\u003C\/p\u003E\u003Cp\u003EThey could, for example, cause a major inconvenience if they move important objects like wallets and keys but are unable to tell the owner where they put them.\u003C\/p\u003E\u003Cp\u003EA new benchmark created by Georgia Tech researchers tests the memory capability of vision-language models (VLMs), including Gemini 2.0-flash and GPT-4o, to determine whether they could improve robot performance.\u003C\/p\u003E\u003Cp\u003E\u003Ca href=\u0022https:\/\/faculty.cc.gatech.edu\/~zk15\/\u0022\u003E\u003Cstrong\u003EZsolt Kira\u003C\/strong\u003E\u003C\/a\u003E, an associate professor in Tech\u2019s School of Interactive Computing, said that most active VLMs can process only a few hundred images at a time before forgetting them.\u003C\/p\u003E\u003Cp\u003E\u201cThere wasn\u2019t a benchmark that could test these specialized capabilities that we might want from a robot,\u201d Kira said. \u201cOur purpose in creating one is to spur research in this area so that others can develop methods to solve this problem.\u201d\u003C\/p\u003E\u003Cp\u003EThe team\u2019s benchmark consists of memory evaluations based on 60 tasks that require continuous engagement and contextual and environmental awareness. Three task categories \u2014 spatial, temporal, and multi-goal \u2014 are highly difficult.\u003C\/p\u003E\u003Cp\u003E\u201cYou can chat with a language model like ChatGPT, and it remembers all the text from your previous conversations with it because the text data is small in terms of storage,\u201d Kira said.\u0026nbsp;\u003C\/p\u003E\u003Cp\u003E\u201cEmbodied agents and robots observe a lot of videos, and those tend to be a lot harder to recall. If you ask it where your diary is, it must efficiently store images of your diary and where it is.\u003C\/p\u003E\u003Cp\u003E\u201cThat\u2019s just a few frames out of a million that it recorded since yesterday, and it must know where that is relative to the spatial layout of the house.\u201d\u003C\/p\u003E\u003Cp\u003E\u003Cstrong\u003EKarmesh Yadav\u003C\/strong\u003E, a Ph.D. student in Kira\u2019s lab, said users don\u2019t want to repeat information about their preferences after deployment. They also expect the robot to maintain long-term memory.\u003C\/p\u003E\u003Cp\u003E\u201cIn reality, you would want it to have a memory of the entire time it has been in your house, whether that\u2019s days or years,\u201d Yadav said.\u003C\/p\u003E\u003Cp\u003EThe team\u2019s benchmark evaluates memory performance over an average one-day period. It challenges the models with commands like \u201cnavigate to an object you did not interact with yesterday\u201d or \u201cnavigate to a room you did not visit yesterday.\u201d\u003C\/p\u003E\u003Cp\u003EThe results show that current state-of-the-art models have a long way to go before they can be used in embodied agents. The best model achieved roughly a 50% success rate on high-level difficulty tasks.\u003C\/p\u003E\u003Cp\u003ESurprisingly, this model wasn\u2019t Gemini or GPT VLMs. It was an open-source reasoning model from \u003Ca href=\u0022https:\/\/qwen.ai\/qwenchat\u0022\u003EQwen\u003C\/a\u003E.\u0026nbsp;\u003C\/p\u003E\u003Cp\u003E\u201cIf you use an open-source reasoning model, even one that\u2019s small in size, it can match the performance of closed-source non-reasoning models,\u201d Yadav said. \u201cReasoning models can look at the video and translate it into text. It describes what\u2019s going on in the video in words and determines the right frame to choose in its final output.\u201d\u003C\/p\u003E\u003Cp\u003EKira said that converting images into descriptive text that language models easily understand could be one way to circumvent the memory storage dilemma. Another way could be to assign a hierarchy that teaches the model to identify irrelevant or redundant images and erase them from storage.\u003C\/p\u003E\u003Cp\u003EWhatever future progress researchers make in those areas, Kira and his students are confident that their benchmark will remain relevant.\u0026nbsp;\u003C\/p\u003E\u003Cp\u003E\u201cThis benchmark is scalable,\u201d said Ph.D. student \u003Cstrong\u003EYusuf Ali\u003C\/strong\u003E. \u201cIf two years from now, videos are easy for a model to understand and store, you can use the same infrastructure we propose and scale up the problem\u2019s difficulty.\u201d\u003C\/p\u003E\u003Cp\u003EYadav and Ali are co-first authors of a paper on the benchmark, presented in early September at the 2026 European Conference on Computer Vision in Sweden.\u003C\/p\u003E\u003Cp\u003EFor more information about the project, click\u0026nbsp;\u003Ca href=\u0022https:\/\/findingdory-benchmark.github.io\/\u0022\u003Ehere\u003C\/a\u003E.\u003C\/p\u003E","summary":"","format":"limited_html"}],"field_subtitle":"","field_summary":[{"value":"\u003Cp\u003ETo improve robot performance and advance memory capabilities needed for reliable home use, Georgia Tech researchers have developed a new benchmark to evaluate the memory capabilities of vision-language models (VLMs).\u003C\/p\u003E","format":"limited_html"}],"field_summary_sentence":[{"value":"To improve robot performance and advance memory capabilities needed for reliable home use, Georgia Tech researchers have developed a new benchmark to evaluate the memory capabilities of vision-language models (VLMs)."}],"uid":"36530","created_gmt":"2026-09-24 19:19:33","changed_gmt":"2026-09-24 19:21:36","author":"Nathan Deen","boilerplate_text":"","field_publication":"","field_article_url":"","location":"Atlanta, GA","dateline":{"date":"2026-09-24T00:00:00-04:00","iso_date":"2026-09-24T00:00:00-04:00","tz":"America\/New_York"},"extras":[],"hg_media":{"681247":{"id":"681247","type":"image","title":"2X6A9069.jpg","body":null,"created":"1790277587","gmt_created":"2026-09-24 19:19:47","changed":"1790277587","gmt_changed":"2026-09-24 19:19:47","alt":"Zsolt Kira sitting in office chair surrounded by old computers","file":{"fid":"265616","name":"2X6A9069.jpg","image_path":"\/sites\/default\/files\/2026\/09\/24\/2X6A9069.jpg","image_full_path":"http:\/\/hg.gatech.edu\/\/sites\/default\/files\/2026\/09\/24\/2X6A9069.jpg","mime":"image\/jpeg","size":153574,"path_740":"http:\/\/hg.gatech.edu\/sites\/default\/files\/styles\/740xx_scale\/public\/2026\/09\/24\/2X6A9069.jpg?itok=OK71DyFs"}}},"media_ids":["681247"],"groups":[{"id":"47223","name":"College of Computing"},{"id":"1188","name":"Research Horizons"},{"id":"50876","name":"School of Interactive Computing"}],"categories":[{"id":"153","name":"Computer Science\/Information Technology and Security"},{"id":"152","name":"Robotics"}],"keywords":[],"core_research_areas":[{"id":"193655","name":"Artificial Intelligence at Georgia Tech"},{"id":"39501","name":"People and Technology"},{"id":"39521","name":"Robotics"}],"news_room_topics":[],"event_categories":[],"invited_audience":[],"affiliations":[],"classification":[],"areas_of_expertise":[],"news_and_recent_appearances":[],"phone":[],"contact":[],"email":[],"slides":[],"orientation":[],"userdata":""}}}