Velokey

Mises à jour de l'API

Restez informé des derniers changements et améliorations de toutes nos API.

modelsmigration

New Models | Seedream 5.0 Pro and Grok Composer 2.5 Fast

ByteDance's flagship image model and xAI's high-speed coding model are now available, alongside a naming cleanup for the Gemini image line.

Seedream 5.0 Pro

  • Model ID: doubao-seedream-5-0-pro
  • Unifies text-to-image, image-to-image, and image editing in one model
  • Output: up to 4K ultra-HD, with precise in-image text rendering
  • Reference fusion: multi-reference editing through the standard async image task workflow
  • Billed per image by clarity tier

Grok Composer 2.5 Fast

  • Model ID: grok-composer-2.5-fast
  • xAI's high-speed model for agentic coding: multi-step edits, tool calling, and terminal operations at low latency
  • A fit for coding agents, IDE assistants, and high-throughput automated development

Naming migration | Gemini image models

  • Nano Banana 2 and Nano Banana Pro are the new display names for gemini-3.1-flash-image-preview and gemini-3-pro-image-preview, matching their common names
  • Model IDs are unchanged — existing requests keep working with no action needed

Usage logs

  • Image task logs now record the resolution parameter, and the task detail dialog shows it alongside quality
models

New Models | GPT-5.6 and Grok 4.5 Now Available

Two flagship releases join the catalog, both live on the OpenAI-compatible endpoint with below-official pricing.

GPT-5.6

  • Model ID: gpt-5.6
  • OpenAI's newest flagship, GA on July 9, 2026 — built for agentic coding, long-horizon reasoning, and tool use
  • Native Programmatic Tool Calling and predictable caching; leads agentic benchmarks like Terminal-Bench and the Coding Agent Index
  • The gpt-5.6-sol, gpt-5.6-terra, and gpt-5.6-luna tiers all route through the canonical gpt-5.6 base — pick a tier per request, billing follows the tier you call

Grok 4.5

  • Model ID: grok-4.5
  • xAI's newest flagship, strong at real-time information, complex reasoning, and agentic tool use
  • Native live web and X search, a strong reasoning mode, and function calling
  • A fit for up-to-date Q&A, research agents, and automation

Getting started

  • Both models work with existing keys — change the model field and send
  • Reasoning effort suffixes (-low, -medium, -high) are supported on both
console

Console | Overview Redesign, Usage Pagination and Dynamic Top-Up

A round of console upgrades focused on the overview dashboard, usage analysis, and billing.

Overview

  • Redesigned KPI cards with a unified look and a cleaner spend-trend chart
  • Key spending ranking: see which API keys consume the most, at a glance
  • Recent errors: failed requests and failed async tasks in one table, with full error text on hover

Usage

  • The usage detail table now supports pagination and sorting by any metric (requests, tokens, cost)
  • Async task failures are included in success-rate and error statistics

Billing

  • Top-up amounts, payment methods, discounts, and exchange rates are now fetched live — promotions apply without a page release
  • Transaction history hides stale unpaid orders after 30 minutes and displays all amounts in USD

Fixes

  • Balance now refreshes within 5 minutes of a successful top-up
  • Copy-to-clipboard works in non-HTTPS environments
playgroundconsole

Playground | Live Model Playground with Real-Time Cost Estimates

Every model detail page now has a built-in Playground — chat with text models, generate images and video, and see the estimated cost before you run anything.

What you can do

  • Text: streaming chat against any text model, right on its detail page
  • Image and video: full generation workflow with reference-image upload via drag and drop
  • Parameter forms: controls are generated from each model's parameter schema, so every option is valid
  • History: recent generations are kept per model for quick comparison

Real-time cost estimation

  • A price estimate appears before generation and updates as you change parameters
  • Estimates account for resolution tiers, duration, and token-billed image models
  • The estimate endpoint is public API too: POST /pg/estimate

Notes

  • Playground runs use your account balance at the same rates as API calls
  • Playground sessions use scoped tokens that never appear in your API key list
apireliability

API Update | Model Parameter Schemas and Pre-Submit Validation

The model detail API now returns a parameter schema for every media model, and video submissions are validated against it before anything is sent upstream.

Parameter schemas

  • GET /models/:model_name includes a parameters block describing each accepted field: type, enum values, ranges, and defaults
  • Schemas cover resolution and aspect-ratio options, duration limits, reference-image counts, and mode flags
  • The Playground and console build their parameter forms directly from this schema, so what you see always matches what the API accepts

Pre-submit validation

  • Invalid parameter combinations are rejected immediately with a specific error message — no more opaque service_busy responses for local validation failures
  • Model-not-found and pricing-not-configured errors now return model_not_found instead of a generic error

Reliability

  • Graceful shutdown: in-flight async tasks survive deploys
  • Stuck image tasks are detected and auto-recovered on startup
models

New Models | Claude Fable 5 and Claude Sonnet 5 Now Available

Anthropic's newest generation arrives on Velokey. Both models are live behind the OpenAI-compatible endpoint — swap the model name and go.

Claude Fable 5

  • Model ID: claude-fable-5
  • Anthropic's Mythos-class frontier model, state-of-the-art across nearly every tested capability
  • Strengths: agentic software engineering, vision processing, long-horizon memory
  • Billed per token at below-list pricing, with cached-token discounts

Claude Sonnet 5

  • Model ID: claude-sonnet-5
  • A substantial upgrade over Sonnet 4.6 with performance approaching Opus 4.8 at a lower cost
  • Context window: 1M tokens, max output: 64K tokens
  • A fit for autonomous multi-step execution, computer use, and cost-efficient agentic runs

Also added this week

  • gpt-5.4-mini — OpenAI's compact tier for high-throughput workloads
  • sora-2 — OpenAI video generation via the async task API
  • wan2.7-image — Alibaba's image line joins the Wan family
models

New Models | Kling v3, Vidu Q3, PixVerse, HappyHorse 1.1, Wan 2.7 and Grok Imagine Video

Six video model families land on Velokey at once — all available through the unified async task API with multi-channel failover.

What's new

  • Kling v3: kling-v3-video-generation and kling-v3-omni-video-generation — text-to-video, image-to-video, reference and edit modes with keep_original_sound support
  • Vidu Q3: viduq3-pro and viduq3-turbo lines covering text-to-video, image-to-video, and start/end-frame-to-video
  • PixVerse: pixverse-v6 and pixverse-c1 families with text, image, reference, and keyframe modes
  • HappyHorse 1.1: happyhorse-1.1-t2v, happyhorse-1.1-i2v, happyhorse-1.1-r2v — watermark-free output by default
  • Wan 2.7: wan2.7-t2v, wan2.7-i2v, wan2.7-r2v, plus wan2.7-videoedit for video editing
  • Grok Imagine Video: grok-imagine-video — xAI's video generation, async task workflow

Notes

  • Mode detection is automatic: sending reference images to a general model ID selects the right sub-mode (for example, more than 2 images upgrades to reference-to-video)
  • Duration values accept both integer and float forms (10 and 10.0)
modelspricing

Model Update | gpt-image-2 Now Billed by Actual Tokens

Token-based billing comes to gpt-image-2 — flat per-image pricing is replaced, and async image generations now settle against the real token usage reported by the model.

What changed

  • Billing basis: each task is charged from the model's reported input and output tokens instead of a fixed per-image rate
  • Pre-hold and settle: an estimated amount is held at submission and reconciled to actual usage when the task completes
  • Resolution-aware estimates: the pre-hold takes the requested resolution tier into account, so the estimate tracks the final cost closely

Why this matters

  • Small or simple generations get cheaper — you no longer pay the worst-case flat rate
  • Cost scales transparently with output size and complexity, matching OpenAI's own billing model

Other image models keep their existing tier or per-call pricing.

pricingmodels

Pricing | Tiered Media Pricing by Resolution and Duration

Image and video models now use tiered pricing that reflects what you actually generate: clarity tiers for images, resolution × duration for video, and per-call pricing where providers bill flat rates.

Image models

  • Output tiers such as 1K, 2K, and 4K are billed independently
  • The model market shows the full tier table for every image model, in credits with USD equivalents

Video models

  • Price is determined by resolution tier and clip duration in seconds
  • Tier tables appear on each model's detail page, so cost is predictable before you submit — for example doubao-seedream-4-5 image tiers or kling-v3-video-generation duration tiers

Per-call models

  • Models with flat provider pricing are billed per successful task, shown as a single-row tier table

Existing requests keep working unchanged; only the pricing display and billing granularity are new.

reliabilityapi

Reliability | Multi-Channel Failover for Video Generation Tasks

Video tasks now retry across multiple upstream channels automatically — at both the submit and the polling stage — so a single provider outage no longer fails your request.

How failover works

  • Submit stage: if the first channel rejects or errors, the task is retried on the next available channel, ordered by price
  • Polling stage: tasks that stall or fail upstream mid-generation are resubmitted to an alternative channel when one exists
  • Exhaustive mode: retries sweep every priority tier round by round instead of stopping at the first tier
  • Billing: you are charged once, only for the attempt that succeeds

Scope

  • Applies to all async video models, including kling-v3-video-generation and kling-v3-omni-video-generation
  • Text relay requests gain the same exhaustive retry mode where multiple channels serve one model
  • No client changes required; the task_id you hold stays valid across internal retries
pricing

Pricing | Calls Are Now Billed at Your Lowest-Rate Group Automatically

Billing now always applies the lowest rate group your account qualifies for — the price you see in the model market is the price you pay, with no manual group switching.

What changed

  • Automatic selection: every call is billed using the cheapest group available to your account, per model
  • Market alignment: discounted prices shown on model pages and the pricing page are the effective billed rates
  • Savings metric: the console dashboard's "saved vs. official" figure is now computed from your actual group rate instead of a flat multiplier

Details

  • Group discounts stack with cached-token pricing for supported text models
  • No API changes are required — existing keys pick up the behavior automatically
apimodels

API Update | Reasoning Effort Suffixes for Text Models

You can now control reasoning depth by appending an effort suffix to any reasoning-capable model name — no extra request parameters needed.

Usage

  • Append -low, -medium, or -high to the model name, for example claude-opus-4-8-high or qwen3.7-max-low
  • The suffix is stripped before channel matching, so routing, availability, and pricing follow the base model
  • Requests without a suffix keep each model's default effort level

Supported models

  • Anthropic: claude-opus-4-8 and the Claude 4 reasoning line
  • Alibaba: qwen3.7-max, qwen3.7-plus
  • xAI: grok-4.3 reasoning variants
  • Any future reasoning model is supported automatically — the suffix convention is applied platform-wide

Compatibility

  • Explicit reasoning_effort in the request body still works and takes precedence over the suffix
console

Console | Request-Level Usage Logs with Cost Breakdown

The console's Usage page now records every API call with request-level detail, so you can trace exactly where your credits go.

What's included

  • Per-request rows: model, endpoint, latency, HTTP status, and the API key that made the call
  • Token breakdown: prompt tokens, completion tokens, and cached tokens are itemized separately for text models
  • Cost column: each row shows the exact credits charged, computed from the live pricing table at call time
  • Filtering: narrow logs by model, API key, status, or time range
  • Task logs: async image and video tasks get their own tab with task status and generation parameters

Notes

  • Logs are retained for 30 days
  • Request and task IDs can be copied with one click for support inquiries
api

Platform Update | Unified Async Task API for Image and Video Generation

Image and video generation on Velokey now runs through a single asynchronous task workflow. Submit a generation request, receive a task ID immediately, and poll for results — the same flow across every media model, from kling-v3-video-generation to gpt-image-2.

How it works

  • Submit: send a standard generation request to the unified media endpoint; the response returns a task_id right away
  • Poll: query the task endpoint with your task_id to track progress
  • Statuses: tasks move through queued, processing, succeeded, or failed
  • Results: completed tasks return output URLs valid for 24 hours — download and persist them on your side

Why async

  • Video generation can take minutes; long-lived HTTP connections are fragile and time out behind most proxies
  • Failed submissions are billed nothing; you are only charged when a task reaches succeeded
  • The same polling contract applies to every provider we route to, so switching models requires no code changes