← Back to Paper List

HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, Y. Zhuang
Zhejiang University, Microsoft Research Asia
Neural Information Processing Systems (2023)
Agent MM Speech Reasoning

📝 Paper Summary

Agentic AI Multi-call tool use with flexible plan Multimodal language models
HuggingGPT uses a large language model as a central controller to plan tasks, dynamically select specialized models from open communities, and synthesize results for complex multimodal requests.
Core Problem
Current large language models lack the ability to process complex multimodal information and cannot autonomously schedule multiple sub-tasks requiring specialized external models.
Why it matters:
  • Language models are restricted to text generation, limiting their utility in vision, speech, and multimodal tasks.
  • Complex real-world tasks require coordinating multiple specialized AI models, which is beyond the capability of standalone language models.
  • Language models often underperform fine-tuned expert models in challenging zero-shot or few-shot domains.
Concrete Example: When asked to 'describe an image and count the objects in it', a standard language model cannot process the image. HuggingGPT decomposes this into image classification, captioning, and object detection, dynamically routing to specialized vision models to synthesize the answer.
Key Novelty
Language-interfaced Large Language Model (LLM) Controller for Open Model Communities
  • Treats the language model as a brain that parses user requests into a dependent graph of sub-tasks.
  • Uses natural language model descriptions as a generic interface to dynamically select expert models from the Hugging Face community.
  • Executes the selected models and aggregates their structured outputs into a final human-readable response.
Architecture
Architecture Figure Figure 1 and 2 (Text-based conceptual flowchart)
The 4-stage workflow of HuggingGPT orchestrating AI models to fulfill a user request.
Breakthrough Assessment
8/10
Pioneering work demonstrating how large language models can act as central controllers to dynamically orchestrate open-source specialized models via natural language descriptions.
×