RAG Evaluation
The RAG Evaluation feature allows users to systematically test and benchmark the retrieval quality of their knowledge collections, helping ensure that RAG (Retrieval-Augmented Generation) pipelines return accurate and relevant results.
Key Capabilities
Dashboard Overview A summary view showing key metrics at a glance.
- Datasets - Set/edit the agent's display name.
- Ready - Datasets ready to be run.
- Runs - Total number of evaluation runs executed
- Search - Quickly find existing datasets by name.
Dataset Creation Users can generate a new evaluation dataset by configuring:
- Knowledge Collection - select which knowledge collection the dataset will be generated from
- Dataset Name - a descriptive name (e.g. "Q2 Retrieval Baseline")
- Description - (optional), notes on what the dataset is evaluating
- Size Tier - choose the number of generated questions (e.g. Standard - 100 questions), with a trade-off noted between coverage and generation time