DARS: Deep AI Research Systems—sophisticated agentic systems capable of dynamic reasoning, adaptive planning, and multi-iteration web retrieval to autonomously conduct complex research workflows
ASI: Artificial Superintelligence—a theoretical AI system that surpasses human intelligence, envisioned here as the end goal of AI recursive self-improvement
LLM-as-a-judge: Using a Large Language Model to automatically evaluate the quality of text generated by another AI system
Faithfulness score: A metric measuring the proportion of cited claims that are actually supported by their referenced sources (accuracy of citations)
Groundedness score: A metric measuring the proportion of all factual claims in the text that have explicit citation support (citation coverage)
Coverage score: A metric calculating the weighted sum of expert-designed rubric criteria met by the generated research response
Rubric Assessment: An evaluation method using expert-designed, weighted criteria to assess conceptual depth, theoretical understanding, and insight quality
Factual Assessment: An automated evaluation framework that verifies citation accuracy and content groundedness using an extract model and a judge model
URL-claim-context triplet: A data structure containing a factual claim extracted from the response, its surrounding context, and its corresponding citation URL for verification