Run repeatable model and agent eval suites and inspect scoring traces with Inspect AI

Run benchmark-style eval suites against models or agents, then inspect scored traces instead of relying on ad hoc chats and gut feel.

Run repeatable model and agent eval suites and inspect scoring traces with Inspect AI

Run benchmark-style eval suites against models or agents, then inspect scored traces instead of relying on ad hoc chats and gut feel.

Installation

Method 1, Agent Skill Exchange

Method 2, Git clone

git clone https://github.com/agentskillexchange/skills.git && cd skills/skills/run-repeatable-model-and-agent-eval-suites-and-inspect-scoring-traces-with-inspect-ai

Method 3, Download ZIP

  • Download the repository ZIP and extract skills/run-repeatable-model-and-agent-eval-suites-and-inspect-scoring-traces-with-inspect-ai.

Method 4, Manual copy

  • Copy this skill folder into your local skills directory, then reload your agent tooling.

Method 5, Fork and sync

  • Fork the repository if you want to maintain local edits while syncing upstream changes.

Source