RAG: Retrieval-Augmented Generation—AI systems that answer questions by first searching for relevant documents.
Instruction Simulation: A technique that uses questions from existing high-quality synthetic datasets as exemplars to guide the format, style, and difficulty of newly generated RAG instructions.
RAG paradigms: Five defined query-document relationships (Single-Doc Answer, Single-Doc Support, Useless Doc, Multi-Doc Answer, Multi-Doc Support) that categorize how helpful retrieved documents are.
GPT-4o: A highly capable proprietary large language model from OpenAI used here to synthesize the instruction data.
Contriever: A pre-trained dense retrieval model used to fetch relevant and unrelated documents during dataset construction and inference.
SFT: Supervised Fine-Tuning—training a model on labeled examples to follow instructions.
vLLM: A high-throughput and memory-efficient LLM inference engine.
LLMs: Large Language Models.