Spaces:
Runtime error
Runtime error
| # RAG Evaluation Module | |
| This folder contains a lightweight retrieval-evaluation harness for the project. | |
| ## Supported Steps | |
| 1. `beir/scifact` | |
| 2. `beir/fiqa` | |
| 3. `open-ragbench` | |
| 4. `t2-ragbench` | |
| 5. `local-options` | |
| Each run builds a temporary Chroma index under `eval/indexes/` and writes reports under `eval/reports/`. | |
| ## Smoke Tests | |
| ```bash | |
| uv --cache-dir .uv-cache run python -m eval.rag_eval --dataset beir/scifact --max-corpus-docs 200 --max-queries 10 --rebuild | |
| uv --cache-dir .uv-cache run python -m eval.rag_eval --dataset beir/fiqa --max-corpus-docs 500 --max-queries 10 --rebuild | |
| uv --cache-dir .uv-cache run python -m eval.rag_eval --dataset open-ragbench --max-corpus-docs 50 --max-queries 10 --rebuild | |
| uv --cache-dir .uv-cache run python -m eval.rag_eval --dataset t2-ragbench --max-corpus-docs 50 --max-queries 10 --rebuild | |
| uv --cache-dir .uv-cache run python -m eval.rag_eval --dataset local-options --max-queries 3 --rebuild | |
| ``` | |
| ## Run The Whole Suite | |
| ```bash | |
| uv --cache-dir .uv-cache run python -m eval.run_eval_suite --rebuild | |
| ``` | |
| By default, the suite runs: | |
| - `beir/scifact` | |
| - `beir/fiqa` | |
| - `open-ragbench` | |
| - `local-options` | |
| Useful options: | |
| ```bash | |
| # Accurate run after changing PDF parsing, chunking, embedding, retrieval code, or sampling parameters. | |
| uv --cache-dir .uv-cache run python -m eval.run_eval_suite --rebuild | |
| # Faster run that reuses existing indexes. | |
| uv --cache-dir .uv-cache run python -m eval.run_eval_suite | |
| # Run only selected datasets. | |
| uv --cache-dir .uv-cache run python -m eval.run_eval_suite --datasets local-options,beir/fiqa | |
| # Override shared parameters for all selected datasets. | |
| uv --cache-dir .uv-cache run python -m eval.run_eval_suite --top-k 10 --max-queries 20 --max-corpus-docs 1000 | |
| # Save a stable suite-level report name. | |
| uv --cache-dir .uv-cache run python -m eval.run_eval_suite --output-name latest_rag_eval | |
| ``` | |
| The suite writes per-dataset reports and one aggregate report under `eval/reports/`. | |
| ## Common Commands | |
| Run with the default multilingual embedding model: | |
| ```bash | |
| uv --cache-dir .uv-cache run python -m eval.run_eval_suite --rebuild | |
| ``` | |
| Use a custom embedding model for experiments: | |
| ```bash | |
| RAG_EMBED_MODEL=intfloat/multilingual-e5-base \ | |
| uv --cache-dir .uv-cache run python -m eval.run_eval_suite \ | |
| --datasets local-options \ | |
| --top-k 5 \ | |
| --output-name local_options_e5_base \ | |
| --rebuild | |
| ``` | |
| Run the fastest local check while developing PDF parsing or chunking: | |
| ```bash | |
| uv --cache-dir .uv-cache run python -m eval.run_eval_suite \ | |
| --datasets local-options \ | |
| --max-queries 3 \ | |
| --top-k 5 \ | |
| --rebuild | |
| ``` | |
| Run only the standard public retrieval smoke tests: | |
| ```bash | |
| uv --cache-dir .uv-cache run python -m eval.run_eval_suite \ | |
| --datasets beir/scifact,beir/fiqa \ | |
| --rebuild | |
| ``` | |
| Run the financial benchmark only: | |
| ```bash | |
| uv --cache-dir .uv-cache run python -m eval.run_eval_suite \ | |
| --datasets beir/fiqa \ | |
| --max-corpus-docs 1000 \ | |
| --max-queries 50 \ | |
| --top-k 5 \ | |
| --rebuild | |
| ``` | |
| Run the PDF-like benchmark only: | |
| ```bash | |
| uv --cache-dir .uv-cache run python -m eval.run_eval_suite \ | |
| --datasets open-ragbench \ | |
| --max-corpus-docs 100 \ | |
| --max-queries 20 \ | |
| --top-k 5 \ | |
| --rebuild | |
| ``` | |
| Compare different `top-k` values: | |
| ```bash | |
| uv --cache-dir .uv-cache run python -m eval.run_eval_suite \ | |
| --datasets local-options \ | |
| --top-k 3 \ | |
| --output-name local_options_top3 \ | |
| --rebuild | |
| uv --cache-dir .uv-cache run python -m eval.run_eval_suite \ | |
| --datasets local-options \ | |
| --top-k 10 \ | |
| --output-name local_options_top10 \ | |
| --rebuild | |
| ``` | |
| Compare retrieval with and without reranker: | |
| ```bash | |
| uv --cache-dir .uv-cache run python -m eval.run_eval_suite \ | |
| --datasets local-options \ | |
| --top-k 5 \ | |
| --output-name local_options_no_reranker \ | |
| --rebuild | |
| uv --cache-dir .uv-cache run python -m eval.run_eval_suite \ | |
| --datasets local-options \ | |
| --top-k 5 \ | |
| --use-reranker \ | |
| --reranker-candidates 25 \ | |
| --output-name local_options_with_reranker \ | |
| --rebuild | |
| ``` | |
| Use a custom reranker model: | |
| ```bash | |
| uv --cache-dir .uv-cache run python -m eval.run_eval_suite \ | |
| --datasets beir/fiqa \ | |
| --use-reranker \ | |
| --reranker-model cross-encoder/ms-marco-MiniLM-L-6-v2 \ | |
| --reranker-candidates 50 \ | |
| --top-k 5 \ | |
| --rebuild | |
| ``` | |
| Compare different chunk settings: | |
| ```bash | |
| uv --cache-dir .uv-cache run python -m eval.run_eval_suite \ | |
| --datasets local-options \ | |
| --chunk-size 384 \ | |
| --chunk-overlap 64 \ | |
| --output-name local_options_chunk384 \ | |
| --rebuild | |
| uv --cache-dir .uv-cache run python -m eval.run_eval_suite \ | |
| --datasets local-options \ | |
| --chunk-size 768 \ | |
| --chunk-overlap 128 \ | |
| --output-name local_options_chunk768 \ | |
| --rebuild | |
| ``` | |
| Run a larger, slower evaluation before reporting results: | |
| ```bash | |
| uv --cache-dir .uv-cache run python -m eval.run_eval_suite \ | |
| --datasets beir/scifact,beir/fiqa,open-ragbench,local-options \ | |
| --max-corpus-docs 2000 \ | |
| --max-queries 100 \ | |
| --top-k 5 \ | |
| --output-name full_rag_eval \ | |
| --rebuild | |
| ``` | |
| Stop immediately when one dataset fails: | |
| ```bash | |
| uv --cache-dir .uv-cache run python -m eval.run_eval_suite \ | |
| --datasets beir/scifact,beir/fiqa,open-ragbench,local-options \ | |
| --fail-fast \ | |
| --rebuild | |
| ``` | |
| Run a single dataset directly without the suite wrapper: | |
| ```bash | |
| uv --cache-dir .uv-cache run python -m eval.rag_eval \ | |
| --dataset local-options \ | |
| --max-queries 3 \ | |
| --top-k 5 \ | |
| --rebuild | |
| ``` | |
| ## Suggested Workflow | |
| 1. During development, run `local-options` with a small query count. | |
| 2. After changing PDF extraction, chunking, embeddings, or retrieval code, add `--rebuild`. | |
| 3. Before comparing two versions, use the same `--datasets`, `--max-queries`, `--max-corpus-docs`, `--top-k`, `--chunk-size`, and `--chunk-overlap`. | |
| 4. Use `--output-name` to save stable report names for before/after comparison. | |
| 5. When testing reranker, compare the same dataset once without `--use-reranker` and once with `--use-reranker`. | |
| ## Metrics | |
| - `hit_at_1` | |
| - `hit_at_3` | |
| - `hit_at_5` | |
| - `hit_at_k` | |
| - `mrr` | |
| - `ndcg_at_k` | |
| The public benchmarks test whether the eval pipeline works on standard datasets. The `local-options` benchmark is the project-specific check for PDF parsing, formula extraction, and section-aware chunking. | |