中文
~/home / TypeScript / GitHub
GitHub · TypeScript

claude-video-vision

claude-video-vision is a skill AI agent for Allow AI to automatically watch and interpret video content, including extracting key frames, recognizing speech, and understanding visual context. Pricing: Free. As of 2026-09-25, KanonAgent records 25 upvotes. First indexed by KanonAgent on 2026-08-08.
claude-video-vision is a plugin for Claude Code that allows AI to analyze videos by extracting frames and interpreting audio, enabling video comprehension similar to human viewing.
claude-video-vision — official preview image
25 upvotes
Tracked by Kanon since Aug 8, 2026
Cooling
status
28/100
momentum · conf 0.92
33d
tracked since 2026-08-22
2026-09-17
last meaningful change
4
evidence records

Status basis: npm_downloads_7d momentum fell to 0 per 3d from 131 (<= 40% of the prior window)

📈 Timelinewhat changed, and when
🔎 Known / Unknown 14 of 23 fields unknown
Machine-callable— unknown —
Open source1 verified 1 quote(s)
Self-hostable— unknown —
Bring your own key— unknown —
Autonomy level2 inferred 1 quote(s)
Pricing model— unknown —
Integrationsmcp,gemini,openai,ffmpeg,whisper inferred 2 quote(s)

“Unknown” means we have not verified it — it is not a “no”. Hard filters never treat unknown as false.

🤖 Agent teardown · skill
Job to be doneAllow AI to automatically watch and interpret video content, including extracting key frames, recognizing speech, and understanding visual context.
AutonomyL2 · tool-calling(evidence: "extracts frames via ffmpeg and processes audio via multiple …")
Who it is forAI researchers and content analysis teams
Prerequisitesopen source
Integrationsmcp · gemini · openai · ffmpeg · whisper
PricingFree
Traction · why it is risingSparked interest in the AI video analysis space with 22 Product Hunt votes.
Why it matters

It extends AI’s capabilities beyond text and images, providing a key building block for video understanding and advancing AI agents’ perceptual abilities.

Evidence quotesverbatim, from the product’s own materials

“public GitHub repository (README fetched)”— structural
“extracts frames via ffmpeg and processes audio via multiple backends”— readme
“processes audio via multiple backends (Gemini API, local Whisper, or OpenAI API)”— readme
“MCP server downloads it with yt-dlp”— readme
TypeScript
Signal source: GitHub
Visit official site →
📛 Official badgefor your site / README
claude-video-vision badge
Building claude-video-vision? Pick a style above — the embed code updates live. Deep color control via URL params: bg= / fg= / accent= (hex). It links back to this page.
Share on X
On mobile tap Share for WeChat / RED (Xiaohongshu) / X; on desktop use Copy text and paste into the app.

FAQ

What is claude-video-vision?

claude-video-vision is a plugin for Claude Code that allows AI to analyze videos by extracting frames and interpreting audio, enabling video comprehension similar to human viewing.

What does claude-video-vision do?

Allow AI to automatically watch and interpret video content, including extracting key frames, recognizing speech, and understanding visual context.

Why does claude-video-vision matter?

It extends AI’s capabilities beyond text and images, providing a key building block for video understanding and advancing AI agents’ perceptual abilities.

How much does claude-video-vision cost?

Free

Is claude-video-vision free?

Yes — claude-video-vision has a free tier. Pricing as stated on its own page: Free

Is claude-video-vision open source?

Yes — claude-video-vision is open source.

What does claude-video-vision integrate with?

mcp,gemini,openai,ffmpeg,whisper

How popular is claude-video-vision?

As tracked by KanonAgent: 25 upvotes (first indexed 2026-08-08).

What are the best claude-video-vision alternatives?

Similar AI agents tracked by KanonAgent: open-design, Agent-Reach, deer-flow, ruflo, career-ops, archify.

claude-video-vision alternatives — similar AI agents

open-designAgent-Reachdeer-flowruflocareer-opsarchify

Where this fits — browse the same shelf

AI Agent Skills & PluginsAI agents for Video generation editingAre there open-source AI agents for Video generation editing?Which AI agents for Video generation editing support MCP?AI agents that work with MCP (Model Context Protocol)AI agents that work with Gemini