🐢 Open-Source Evaluation & Testing library for LLM Agents
-
Updated
Aug 4, 2026 - Python
🐢 Open-Source Evaluation & Testing library for LLM Agents
AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.
Dataset and benchmark for RAG on company internal documents.
RAG evaluation without the need for "golden answers"
中文优先的企业 RAG 知识库:可控解析、治理、切块、混合检索、重排、引用、GraphRAG、评测与 Dify 接入。
Red Teaming python-framework for testing chatbots and GenAI systems.
Open-source toolkit for reliable RAG pipelines: convert PDFs to Markdown, clean documents, inspect chunks, compare chunking strategies, and enrich metadata for LLM applications.
开源 AI 应用评测平台,支持 RAG、AI Agent、多轮对话、LLM-as-Judge、接口评测、评测报告和人工盲测。Open-source AI evaluation platform for RAG, AI Agents, multi-turn conversations, LLM-as-Judge, endpoint evaluation
RAG boilerplate with semantic/propositional chunking, hybrid search (BM25 + dense), LLM reranking, query enhancement agents, CrewAI orchestration, Qdrant vector search, Redis/Mongo sessioning, Celery ingestion pipeline, Gradio UI, and an evaluation suite (Hit-Rate, MRR, hybrid configs).
LLM and agent evaluation for Java & Kotlin. Runs in JUnit and CI. Spring AI, LangChain4j, Koog, Embabel, and any LLM client.
Lightweight RAG provenance middleware. Verifies every claim in an LLM response is grounded in a retrieved source - without an LLM call.
Open source framework for evaluating AI Agents
smallevals — CPU-fast, GPU-blazing fast offline retrieval evaluation for RAG systems with tiny QA models.
The E2E AI testing tool | No ML Overhead
OMK — Observe. Measure. Know. Make every knowledge change in your AI application evidence-backed.
⚡️ The "1-Minute RAG Audit" — Generate QA datasets & evaluate RAG systems in Colab, Jupyter, or CLI. Privacy-first, async, visual reports.
This project aims to compare different Retrieval-Augmented Generation (RAG) frameworks in terms of speed and performance.
Add a description, image, and links to the rag-evaluation topic page so that developers can more easily learn about it.
To associate your repository with the rag-evaluation topic, visit your repo's landing page and select "manage topics."