Skill · Research
benchmarking
Compare models on operator reality, not demo sparkle.
Overview
What this skill does
Benchmark models or agents, compare providers for real work, and track performance across runs.
ClawHub download
Install from ClawHub
This skill is mirrored to ClawHub for install-from-the-page convenience.
The ClawHub package is regenerated by scripts/sync-superada-to-clawhub.mjs on every push to main.
Install contract
Where it lives and how to install it
- Source
- enterprise-crew-skills/benchmarking
- Entrypoint
- SKILL.md
- Installability
- Ready for agent install
- Source spec
- github:henrino3/enterprise-crew-skills/benchmarking
This page is designed to be the thing you send to another agent. It tells the agent where the source of truth lives, whether install is real today, and what command or manual path to follow.
Install command
openclaw skills install github:henrino3/enterprise-crew-skills/benchmarking Prompt to hand another agent
Install the skill described on this page from github:henrino3/enterprise-crew-skills/benchmarking. Review the source first, then run: openclaw skills install github:henrino3/enterprise-crew-skills/benchmarking Setup
Install steps
- Step 1
Review the bundle source.
- Step 2
Install from GitHub.
- Step 3
Use it when you need structured benchmark runs and comparison framing.
What it covers
- Skill contract
- Public README
Good use cases
- Compare providers for a task family
- Run scorecard-style evaluations
- Track model quality over time
Files
Published artifacts
- Skill contract
github:henrino3/enterprise-crew-skills/benchmarking/SKILL.mdInstallable benchmarking instructions.
- README
github:henrino3/enterprise-crew-skills/benchmarking/README.mdPublic overview for the benchmark skill.