{"691627":{"#nid":"691627","#data":{"type":"news","title":"New Model Teaches Workplace AI to Read More Like Humans","body":[{"value":"\u003Cp\u003EModern artificial intelligence (AI) platforms can summarize reports, analyze documents, and answer questions in seconds. But when information is spread across dozens of slides, charts, and tables, even advanced models can miss important details.\u0026nbsp;\u003C\/p\u003E\u003Cp\u003EAs many organizations and businesses are turning to AI to improve workplace efficiency, risk remains high. Between overlooked footnotes and misread graphics, small mistakes can have expensive consequences.\u0026nbsp;\u003C\/p\u003E\u003Cp\u003ETo address this challenge, researchers from Georgia Tech and J.P. Morgan developed SlideAgent. The new framework helps large language models (LLMs) better understand complex visual documents like presentation slide decks, brochures, and reports.\u0026nbsp;\u003C\/p\u003E\u003Cp\u003E\u003Ca href=\u0022https:\/\/arxiv.org\/abs\/2510.26615\u0022\u003ESlideAgent\u003C\/a\u003E works by breaking documents into multiple levels, allowing the model to analyze both the big picture and the fine details. This human-inspired approach leads to more accurate and reliable interpretation than existing systems.\u0026nbsp;\u003C\/p\u003E\u003Cp\u003EBeyond improving workplace tools, SlideAgent also points to a broader shift in AI. Instead of building only larger and more powerful models, the work shows how smarter design and more efficient reasoning can improve performance.\u0026nbsp;\u003C\/p\u003E\u003Cp\u003E\u201cMultimodal LLMs such as GPT, Gemini, and Claude, can save people time and reduce the mental effort required to understand these documents, but they remain imperfect,\u201d said \u003Ca href=\u0022https:\/\/ahren09.github.io\/\u0022\u003EYiqiao (Ahren) Jin\u003C\/a\u003E, a Ph.D. candidate in Georgia Tech\u2019s \u003Ca href=\u0022https:\/\/cse.gatech.edu\/\u0022\u003ESchool of Computational Science and Engineering\u003C\/a\u003E (CSE).\u003C\/p\u003E\u003Cp\u003E\u201cIn high-stakes fields such as finance, for example, misreading a number, overlooking a footnote, or making an incorrect comparison across pages could affect reporting, risk assessment, or strategic decisions.\u201d\u0026nbsp;\u003C\/p\u003E\u003Cp\u003EThe researchers tested SlideAgent on a wide range of real-world documents, including financial presentations, technical slides, and visual question-answering datasets.\u0026nbsp;\u003C\/p\u003E\u003Cp\u003EThe system consistently outperformed leading commercial models and open-source tools throughout the evaluation. In some cases, it improved accuracy by up to 10%. The gains were especially strong on more complex tasks, such as comparing information across slides or understanding how visuals relate to each other on a page.\u003C\/p\u003E\u003Cp\u003E\u201cWe found the results very encouraging. SlideAgent reaches an improvement of 7.9% over its proprietary base model and 9.8% over the evaluated open-source base models,\u201d said Jin, the project\u2019s lead researcher.\u0026nbsp;\u003C\/p\u003E\u003Cp\u003E\u201cThese are meaningful gains given the strength of the underlying multimodal models and the difficulty of the tasks.\u201d\u0026nbsp;\u003C\/p\u003E\u003Cp\u003ECurrent multimodal AI systems often process entire pages at once. This approach can lead to mistakes, such as miscounting items in a chart or overlooking important details in dense visuals.\u003C\/p\u003E\u003Cp\u003ESlideAgent addresses this by mimicking how people read documents. Instead of treating each page as a single unit, the system looks at information at three levels: the full document, individual pages, and specific elements like charts, tables, and text blocks.\u003C\/p\u003E\u003Cp\u003EA network of agents, each specialized for a specific level, divides and coordinates analysis. Then, SlideAgent combines outputs to build a structured understanding of the overall document. This allows it to answer questions more accurately and reason across multiple pages.\u003C\/p\u003E\u003Cp\u003E\u201cThe central inspiration was how people naturally read a long presentation,\u201d Jin said.\u003C\/p\u003E\u003Cp\u003E\u201cWe first develop an understanding of the overall narrative, then identify the relevant pages or sections, and finally zoom in on individual charts, tables, or text blocks when precise evidence is needed.\u201d\u003C\/p\u003E\u003Cp\u003EThe work highlights a growing challenge with AI. As systems become more popular and more powerful, users increasingly discover the technology\u2019s limitations. This is especially true for real-world tasks that require structured reasoning and contextual understanding.\u003C\/p\u003E\u003Cp\u003ESlideAgent shows that better performance does not always come from building bigger models. Instead, it points to smarter ways of organizing how AI processes information that can lead to improvement.\u003C\/p\u003E\u003Cp\u003EThe Association for Computational Linguistics (ACL) accepted SlideAgent for presentation at its annual meeting. The 64th \u003Ca href=\u0022https:\/\/2026.aclweb.org\/\u0022\u003EACL 2026 meeting\u003C\/a\u003E took place July 2-7 in San Diego.\u003C\/p\u003E\u003Cp\u003EACL is a scientific and professional organization for researchers in natural language processing (NLP). Its namesake conference is one of the world\u2019s leading venues for presenting NLP research.\u003C\/p\u003E\u003Cp\u003EACL 2026 followed a year after Jin completed an internship at J.P. Morgan AI Research, where he worked on SlideAgent. He interned under \u003Cstrong\u003ERachneet Kaur\u003C\/strong\u003E, \u003Cstrong\u003EZhen Zeng\u003C\/strong\u003E, and \u003Cstrong\u003ESumitra Ganesh\u003C\/strong\u003E, all co-authors of the paper. School of CSE Associate Professor \u003Ca href=\u0022https:\/\/faculty.cc.gatech.edu\/~srijan\/\u0022\u003ESrijan Kumar\u003C\/a\u003E advises Jin at Georgia Tech.\u003C\/p\u003E\u003Cp\u003EAlong with SlideAgent, Jin authored two other papers accepted at ACL 2026.\u003C\/p\u003E\u003Cp\u003E\u201cConferences such as ACL are valuable not only for sharing results, but also for refining future project ideas with the broader community. Discussions across institutions and research areas can reveal limitations, suggest new evaluations, and spark collaborations that are difficult to develop in isolation,\u201d Jin said.\u003C\/p\u003E\u003Cp\u003E\u201cFor SlideAgent, I was particularly interested in feedback on how we can make the framework more computationally efficient without sacrificing accuracy or interpretability.\u201d\u003C\/p\u003E","summary":"","format":"limited_html"}],"field_subtitle":"","field_summary":[{"value":"\u003Cp\u003EModern artificial intelligence (AI) platforms can summarize reports, analyze documents, and answer questions in seconds. But when information is spread across dozens of slides, charts, and tables, even advanced models can miss important details.\u0026nbsp;\u003C\/p\u003E\u003Cp\u003EAs many organizations and businesses are turning to AI to improve workplace efficiency, risk remains high. Between overlooked footnotes and misread graphics, small mistakes can have expensive consequences.\u0026nbsp;\u003C\/p\u003E\u003Cp\u003ETo address this challenge, researchers from Georgia Tech and J.P. Morgan developed SlideAgent. The new framework helps large language models (LLMs) better understand complex visual documents like presentation slide decks, brochures, and reports.\u0026nbsp;\u003C\/p\u003E\u003Cp\u003ESlideAgent works by breaking documents into multiple levels, allowing the model to analyze both the big picture and the fine details. This human-inspired approach leads to more accurate and reliable interpretation than existing systems.\u0026nbsp;\u003C\/p\u003E\u003Cp\u003EBeyond improving workplace tools, SlideAgent also points to a broader shift in AI. Instead of building only larger and more powerful models, the work shows how smarter design and more efficient reasoning can improve performance.\u0026nbsp;\u003C\/p\u003E","format":"limited_html"}],"field_summary_sentence":[{"value":"The new AI framework called SlideAgent helps large language models (LLMs) better understand complex visual documents used in the workplace like presentation slide decks, brochures, and reports. "}],"uid":"36319","created_gmt":"2026-08-12 12:01:25","changed_gmt":"2026-08-12 12:09:10","author":"Bryant Wine","boilerplate_text":"","field_publication":"","field_article_url":"","location":"Atlanta, GA","dateline":{"date":"2026-08-12T00:00:00-04:00","iso_date":"2026-08-12T00:00:00-04:00","tz":"America\/New_York"},"extras":[],"hg_media":{"680843":{"id":"680843","type":"image","title":"SlideAgent-Head-Image.jpg","body":null,"created":"1786536102","gmt_created":"2026-08-12 12:01:42","changed":"1786536102","gmt_changed":"2026-08-12 12:01:42","alt":"Yiqiao (Ahren) Jin ACL 2026","file":{"fid":"265169","name":"SlideAgent-Head-Image.jpg","image_path":"\/sites\/default\/files\/2026\/08\/12\/SlideAgent-Head-Image.jpg","image_full_path":"http:\/\/hg.gatech.edu\/\/sites\/default\/files\/2026\/08\/12\/SlideAgent-Head-Image.jpg","mime":"image\/jpeg","size":96345,"path_740":"http:\/\/hg.gatech.edu\/sites\/default\/files\/styles\/740xx_scale\/public\/2026\/08\/12\/SlideAgent-Head-Image.jpg?itok=MXXpRJcQ"}},"680844":{"id":"680844","type":"image","title":"YJ_ACL2026.jpg","body":null,"created":"1786536492","gmt_created":"2026-08-12 12:08:12","changed":"1786536492","gmt_changed":"2026-08-12 12:08:12","alt":"Yiqiao (Ahren) Jin ACL 2026","file":{"fid":"265170","name":"YJ_ACL2026.jpg","image_path":"\/sites\/default\/files\/2026\/08\/12\/YJ_ACL2026.jpg","image_full_path":"http:\/\/hg.gatech.edu\/\/sites\/default\/files\/2026\/08\/12\/YJ_ACL2026.jpg","mime":"image\/jpeg","size":48282,"path_740":"http:\/\/hg.gatech.edu\/sites\/default\/files\/styles\/740xx_scale\/public\/2026\/08\/12\/YJ_ACL2026.jpg?itok=Vns4uEGI"}},"680845":{"id":"680845","type":"image","title":"YJ_ACL2026.jpeg","body":null,"created":"1786536520","gmt_created":"2026-08-12 12:08:40","changed":"1786536520","gmt_changed":"2026-08-12 12:08:40","alt":"Yiqiao (Ahren) Jin ACL 2026","file":{"fid":"265171","name":"YJ_ACL2026.jpeg","image_path":"\/sites\/default\/files\/2026\/08\/12\/YJ_ACL2026.jpeg","image_full_path":"http:\/\/hg.gatech.edu\/\/sites\/default\/files\/2026\/08\/12\/YJ_ACL2026.jpeg","mime":"image\/jpeg","size":219002,"path_740":"http:\/\/hg.gatech.edu\/sites\/default\/files\/styles\/740xx_scale\/public\/2026\/08\/12\/YJ_ACL2026.jpeg?itok=inGoNwsP"}}},"media_ids":["680843","680844","680845"],"groups":[{"id":"47223","name":"College of Computing"},{"id":"1188","name":"Research Horizons"},{"id":"50877","name":"School of Computational Science and Engineering"}],"categories":[{"id":"194606","name":"Artificial Intelligence"},{"id":"139","name":"Business"},{"id":"153","name":"Computer Science\/Information Technology and Security"},{"id":"194609","name":"Industry"},{"id":"135","name":"Research"},{"id":"134","name":"Student and Faculty"},{"id":"8862","name":"Student Research"}],"keywords":[{"id":"654","name":"College of Computing"},{"id":"166983","name":"School of Computational Science and Engineering"},{"id":"187915","name":"go-researchnews"},{"id":"9153","name":"Research Horizons"},{"id":"10199","name":"Daily Digest"},{"id":"181991","name":"Georgia Tech News Center"},{"id":"170447","name":"Institute for Data Engineering and Science"},{"id":"176858","name":"machine learning center"},{"id":"9167","name":"machine learning"},{"id":"187812","name":"artificial intelligence (AI)"},{"id":"14646","name":"human-computer interaction"},{"id":"192863","name":"go-ai"},{"id":"194384","name":"Tech AI"}],"core_research_areas":[{"id":"193655","name":"Artificial Intelligence at Georgia Tech"},{"id":"39431","name":"Data Engineering and Science"}],"news_room_topics":[],"event_categories":[],"invited_audience":[],"affiliations":[],"classification":[],"areas_of_expertise":[],"news_and_recent_appearances":[],"phone":[],"contact":[{"value":"\u003Cp\u003EBryant Wine, Communications Officer\u003Cbr\u003E\u003Ca href=\u0022mailto:bryant.wine@cc.gatech.edu\u0022\u003Ebryant.wine@cc.gatech.edu\u003C\/a\u003E\u003C\/p\u003E","format":"limited_html"}],"email":[],"slides":[],"orientation":[],"userdata":""}}}