Eval framework. Define correct, test against it, get results.
-
Updated
Feb 17, 2026 - Go
Eval framework. Define correct, test against it, get results.
Find what your AI agent gets wrong — before you have a rubric. Qualitative eval for PMs.
RAG context precision and context recall evaluator measuring retrieved chunk relevance and rank quality
RAG context precision and context recall evaluator measuring retrieved chunk relevance and rank quality
A web-based interactive demo for the GuessArena evaluation framework
One-stop CLI for running structured skill evals across all agent harnesses
Curated AI agent evaluation skills from Microsoft's Eval Guide — plan, generate, run, and interpret eval suites for Copilot Studio agents
4-model parallel planning workflow with eval framework — Claude, Gemini, Codex, GLM-5 · OpenClaw ecosystem
Evaluate and trace tool-calling LLM agents, with one guarantee: a tool failure can never look like an empty result.
Evaluation framework for testing LLM outputs locally. Define prompt templates and custom scorers, run evals against OpenAI or Anthropic models, store results in SQLite, and browse via Rich CLI dashboard or FastAPI UI. Lightweight, self-contained, extensible Python library.
Binary safety verdicts (SAFE/HELD/LEAK/MISS/BROKE) + persona fan-out for LLM pipeline evals
Observability layer for multi-step AI pipelines — traces execution, auto-diagnoses root causes, and builds a growing eval dataset from human feedback
Python tool for extracting validated, schema-defined JSON from documents (PDF, HTML, Markdown, plaintext) using LLMs, with source grounding and a built-in eval harness. Works with Anthropic Claude or OpenAI.
Self-hosted evaluation framework for tool-calling LLM agents. Define tasks in YAML with expected tool calls and scoring rubrics, run against Claude or GPT-4, get scored reports via CLI and FastAPI dashboard. Lightweight, local-first, no lock-in.
🚀 基于Java的开源AI自动化评测框架 / An open source AI automation evaluation framework based on Java
Open-source evaluation framework for LLM agents. Run head-to-head A/B tests, score with LLM-as-judge rubrics, and visualize results in a Streamlit dashboard. Model-agnostic, self-hosted, zero external infrastructure.
Open-source evaluation framework for MCP servers powered by Claude. Auto-discovers tools, generates test scenarios, runs LLM-as-judge scoring to assess correctness and safety, audits for security issues, and visualizes traces in a web dashboard.
Lightweight CLI for versioning prompts and running eval suites. Score outputs with deterministic matching or LLM-as-judge, compare prompt versions with rich terminal diffs. No infra, git-friendly, local-first.
Self-hosted evaluation framework for MCP servers. Define YAML test suites, run agents against MCP tools, score with LLM-as-judge rubrics, and monitor results in a FastAPI dashboard with tool-call traces and regression detection.
Open-source evaluation framework for AI agents. Define test suites with rubrics, run your agent, get LLM-as-judge scores against criteria, inspect full execution traces, and diff runs to catch behavioral regressions.
To associate your repository with the eval-framework topic, visit your repo's landing page and select "manage topics."