Author test
jev-rerank-bench
Test Jev as a search reranker
How this project uses Jev
- Input
- Candidate passages returned by BM25
- Jev decides
- Which passages are relevant and their order
- Code executes
- Evaluation code computes retrieval metrics and exposes raw output
Evidence and limitations
The repository compares multiple datasets. Different averaging methods can change rankings; small average differences do not establish an overall winner.
This project has not been run independently here. Author-reported results are not independently verified results.