ai-evals

Datasets, benchmarks, and evaluation metrics for AI agents. Name benchmark repositories eval-<agent-name>-benchmark.

This public repository is a top-level area of the IT-One organization.

Collaboration

  • Use short-lived branches and pull requests.
  • Keep credentials out of Git; use organization-level Gitea Actions secrets.
  • Add a focused README to every project repository created in this area.
  • Use the AI agent template for agent projects.
S
Description
Datasets, benchmarks, and quality metrics for agents.
Readme
25 KiB