← Back to Paper List

KARMA: Augmenting Embodied AI Agents with Long-and-Short Term Memory Systems

Zixuan Wang, Bo Yu, Jun Zhao, Wenhao Sun, Sai Hou, Shuai Liang, Xingyuan Hu, Yinhe Han, Yiming Gan
Institute of Automation, Chinese Academy of Sciences, Shenzhen Institute of Artificial Intelligence and Robotics for Society, Alibaba Group, Institute of Computing Technology, Chinese Academy of Sciences, Beijing Institute of Technology, University of Chinese Academy of Sciences
IEEE International Conference on Robotics and Automation (2024)
Memory Agent MM

📝 Paper Summary

Agentic RAG pipeline Layered memory Multi-task planning
KARMA equips embodied AI agents with a dual-memory system—a 3D scene graph for long-term environment maps and an adaptive short-term cache for object states—to improve long-sequence task planning.
Core Problem
Embodied AI agents executing long-sequence household tasks often exceed the in-context memory limits of Large Language Models (LLMs), causing them to forget critical object locations and states.
Why it matters:
  • Relying solely on LLM context windows leads to repetitive exploration and redundant actions, drastically reducing task success rates and efficiency in complex environments.
  • Prior memory mechanisms often save everything permanently (causing storage bloat) or use naive first-in-first-out eviction (discarding crucial context).
Concrete Example: When asked to 'wash an apple and place it in a bowl' and later 'bring an apple', an agent without short-term memory forgets the apple's location and state, forcing it to blindly re-explore the kitchen.
Key Novelty
Long-and-Short Term Memory System for Embodied Agents
  • Long-term memory maintains a non-volatile 3D map of the static environment, allowing the agent to understand spatial relationships without re-exploration.
  • Short-term memory acts as a dynamic cache for recently used objects, using an adaptive replacement policy (W-TinyLFU) to keep the most useful information readily available.
Evaluation Highlights
  • Improves Success Rate (SR) by 1.3x (43% vs 33%) and reduces execution time by 68.7% on Composite Tasks compared to CAPEAM.
  • Increases Success Rate (SR) by 2.3x (21% vs 9%) and reduces execution time by 69% on Complex Tasks compared to HELPER.
  • The W-TinyLFU memory replacement policy consistently achieves higher Memory Hit Rates (MHR) than standard FIFO across long task sequences.
Breakthrough Assessment
7/10
A practical and effective architectural design adapting classic computing memory hierarchy (LFU cache + persistent storage) to LLM-based embodied agents, showing strong empirical gains in simulation.
×