LLM: Large Language Model—a powerful AI system trained on massive text to understand and generate human language.
Hugging Face: A public machine learning community and hub hosting numerous pre-trained models for various specific tasks.
JSON: JavaScript Object Notation—a lightweight format for storing and transporting data, used here for structured task outputs.
Task planning: The process where the LLM analyzes a user request and decomposes it into a collection of structured, manageable tasks with execution dependencies.
Model selection: The dynamic assignment of the most appropriate expert model from a repository to a specific sub-task based on the model's text description.
Task execution: Running the selected models with the specified inputs and handling resource dependencies dynamically.
Response generation: Integrating structured outputs from executed models into a final, human-friendly natural language response.
F1: A metric balancing precision and recall.
Edit Distance: A metric measuring the minimum number of operations required to transform one sequence into another.
ViT: Vision Transformer—a model architecture for image processing.
DETR: DEtection TRansformer—an object detection model.
BLIP: Bootstrapping Language-Image Pre-training—a multimodal vision-language model.
ControlNet: A neural network structure to control diffusion models by adding extra conditions.