中文
~/home / Show HN
Show HN

An AI agent skill repo built around evals, not demos

An AI agent skill repo built around evals, not demos is a skill AI agent for Validate agent skills against real performance metrics instead of surface demonstrations. Pricing: not stated on the product page. As of 2026-09-11, KanonAgent records 3 upvotes. First indexed by KanonAgent on 2026-07-24.
The repository supplies AI-agent skills that have been validated through systematic evaluations rather than demo videos.
An AI agent skill repo built around evals, not demos — official preview image
3 upvotes
Tracked by Kanon since Jul 24, 2026
no signal yet
status
16/100
momentum · conf 0.53
20d
tracked since 2026-08-22
6
evidence records

No signal event on file yet — status is unknown, not "quiet".

📈 Timelinewhat changed, and when
🔎 Known / Unknown 13 of 23 fields unknown
Machine-callable unknown
Open source1 inferred 1 quote(s)
Self-hostable unknown
Bring your own key1 inferred 1 quote(s)
Autonomy level2 inferred 2 quote(s)
Pricing model unknown
Integrationsclaude-code,codex inferred 2 quote(s)

“Unknown” means we have not verified it — it is not a “no”. Hard filters never treat unknown as false.

🤖 Agent teardown · skill
Job to be doneValidate agent skills against real performance metrics instead of surface demonstrations.
AutonomyL2 · tool-calling(evidence: "runme eval skills/world-cup-picks-report/evals/regression…")
Who it is forAI researchers and agent developers
Prerequisitesopen source · bring your own API key
Integrationsclaude-code · codex
Traction · why it is risingReceived 3 Product Hunt votes, showing community interest in evaluation standards.
Why it matters

Shifts AI-agent development from demo-driven claims to verifiable engineering practice.

Evidence quotesverbatim, from the product’s own materials

“public GitHub repository (README fetched)”— structural
“Export an Anthropic API key before running the evals:”— readme
“runme eval skills/world-cup-picks-report/evals/regression”— readme
“define your pipeline in YAML”— readme
“/plugin install world-cup-picks-report@sourishkrout-skills”— readme
“--agent codex”— readme
Signal source: Show HN
Visit official site →
📛 Official badgefor your site / README
An AI agent skill repo built around evals, not demos badge
Building An AI agent skill repo built around evals, not demos? Pick a style above — the embed code updates live. Deep color control via URL params: bg= / fg= / accent= (hex). It links back to this page.
Share on X
On mobile tap Share for WeChat / RED (Xiaohongshu) / X; on desktop use Copy text and paste into the app.

FAQ

What is An AI agent skill repo built around evals, not demos?

The repository supplies AI-agent skills that have been validated through systematic evaluations rather than demo videos.

What does An AI agent skill repo built around evals, not demos do?

Validate agent skills against real performance metrics instead of surface demonstrations.

Why does An AI agent skill repo built around evals, not demos matter?

Shifts AI-agent development from demo-driven claims to verifiable engineering practice.

Is An AI agent skill repo built around evals, not demos open source?

Yes — An AI agent skill repo built around evals, not demos is open source.

What does An AI agent skill repo built around evals, not demos integrate with?

claude-code,codex

How popular is An AI agent skill repo built around evals, not demos?

As tracked by KanonAgent: 3 upvotes (first indexed 2026-07-24).

What are the best An AI agent skill repo built around evals, not demos alternatives?

Similar AI agents tracked by KanonAgent: open-design, deer-flow, Agent-Reach, ruflo, career-ops, orca.

An AI agent skill repo built around evals, not demos alternatives — similar AI agents

open-designdeer-flowAgent-Reachruflocareer-opsorca

Where this fits — browse the same shelf

AI Agent Skills & PluginsAI agents for Multi Agent OrchestrationAre there open-source AI agents for Multi Agent Orchestration?AI agents that work with Claude CodeAI agents that work with OpenAI Codex