LLMs: Large Language Models—AI models trained on vast amounts of text to understand and generate human language
VLMs: Vision-Language Models—AI models capable of processing both text and image inputs simultaneously
Deep Research Agents: AI agents capable of complex reasoning, web navigation, file processing, and autonomous knowledge discovery
Agent Foundation Models: Base language or multimodal models specifically pre-trained or fine-tuned to act as the core decision-making engine for AI agents
Persona Hub: A strategy to synthesize large-scale diverse queries by using various personas as triggers for an LLM to generate synthetic tasks
Rejection sampling: A technique where multiple outputs are generated and only those meeting a specific success or similarity criterion are kept for training
GAIA: General AI Assistants—a comprehensive benchmark designed to assess general intelligence and multi-step reasoning capabilities of AI agents
Pass@1: A metric measuring the percentage of times the agent's first attempt successfully solves the task
Pass@3: A metric measuring the percentage of times the agent successfully solves the task within three attempts
SFT: Supervised Fine-Tuning—training a model on labeled examples to follow specific instructions or generate specific formats
Playwright: An open-source automation library used to control web browsers, utilized here by the Web Agent
Accessibility tree: A text-based structural representation of a webpage's elements, used as input for the agent to understand the page without visual processing
Trajectory sampling: The process of recording the step-by-step actions and observations an agent takes to solve a task, used as training data
LangChain: A framework for developing applications powered by language models, used here for similarity-based matching during rejection sampling