Skip to content

Model Pricing

Current list prices for the AI models behind Video Context MCP. All prices are in USD, reflect list prices as of August 2026, and are subject to change by each vendor — always confirm on the vendor pricing pages linked at the bottom of this page.

Prices are per 1M tokens (input / output) unless noted.

ProviderModelInput (per 1M)Output (per 1M)Free tier / notes
Gemini (default)gemini-3.7-flashFreeFreeFree tier, no credit card required. Paid tier: $0.75 input / $3.75 output (→ $1.50 / $7.50 in 2027). Cache $0.075/M.
MiniMaxMiniMax-M3$0.30 (≤512K) / $0.60 (>512K)$1.20 / $2.40Permanent 50% off the standard rate. Cache hit $0.06/M (≤512K). Priority tier = 1.5× standard.
Kimi (Moonshot)kimi-k3$0.60$3.00No free tier.
Alibaba Qwenqwen3.7-plus$0.50$3.00List price; non-thinking requests ≤256K start at $0.40 / $1.60. qwen3.7-flash is the lower-cost variant.
Xiaomi MiMomimo-v2.5$0.40$2.00Free first week (launch).
Z.AI GLMglm-5.3-flash$0.075 (promo)$0.25 (promo)Launch promo ends 2026-09-09, then $0.15 / $0.50. GLM_MODEL=glm-5v-turbo restores the previous model; glm-4.6v-flash keeps the legacy free tier.

Token usage varies with video length and extraction settings, but the project’s live tests give these rough per-call figures:

  • Gemini (default)$0 on the free tier.
  • MiniMax-M3 — roughly $0.02–0.03 per analyze_video / summarize_video call (e.g. ~75K input + ~2.5K output ≈ $0.0255). Follow-up questions on the same video hit the prompt cache and cost ~5× less (≈ $0.005).
  • Qwen3.7 Plus — roughly $0.05 per call.
  • GLM-5.3-Flash — the cheapest paid analysis provider at the launch promo (16× cheaper than GLM-5V-Turbo); still the last-resort fallback because its legacy free tier is aggressively rate-limited. Set GLM_MODEL=glm-4.6v-flash for the free tier.

All media generation tools are Pro-only and billed to your MINIMAX_API_KEY balance. Top up at MiniMax Balance. Source: MiniMax Pay-as-You-Go pricing.

MiniMax-H3 (V2 API, default) is billed per output second:

ResolutionPrice
2K$0.13 / second
768P$0.08 / second

Examples: 5 s @ 2K = $0.65 · 5 s @ 768P = $0.40 · 10 s @ 2K = $1.30. Shorter clips are proportionally cheaper.

Input materials for H3: first 5 reference images free, then $0.04 each; reference video billed per input second at the output resolution rate; reference audio is free.

Related H3 endpoints (not yet exposed as tools): Regeneration (768P→2K) at $0.05/s of regenerated output; H3-Context-IR at $0.90/M input, $3.60/M output.

V1 models are billed per video and are cheaper for casual drafts:

ModelPrice
MiniMax-Hailuo-2.3$0.28 (768P, 6 s) · $0.56 (768P, 10 s) · $0.49 (1080P, 6 s)
MiniMax-Hailuo-2.3-Fast$0.19 (768P, 6 s) · $0.32 (768P, 10 s) · $0.33 (1080P, 6 s)
MiniMax-Hailuo-02$0.10 (512P, 6 s) · $0.15 (512P, 10 s) · $0.28 (768P, 6 s) · $0.56 (768P, 10 s) · $0.49 (1080P, 6 s)
T2V-01 / T2V-01-DirectorBilled per video — see the MiniMax pricing page for current rates
ModelPrice
image-01$0.0035 per image
ModelPrice
music-2.6$0.15 per track (up to 5 minutes)
music-2.6-freeFree (rate-limited, RPM 3)
Lyrics generation$0.01 per song (lyrics_optimizer)

Priced per character:

ModelPrice
speech-2.8-turbo / speech-02-turbo$60 per 1M characters
speech-2.8-hd / speech-02-hd$100 per 1M characters
  • S3 Relay — AWS S3 storage ≈ $0.023/GB/month after the free tier (5 GB + 20K GETs/month for 12 months).
  • Audio transcription — see Audio Providers (Deepgram $200 free credits, AssemblyAI $50, Groq & Gemini free tiers).
  • Pro licensePro Tier costs $10/year (launch promo) and unlocks all media generation tools.

Prices are verified against the project’s source-of-truth repository and vendor pages. Always confirm current rates before relying on them: