Alibaba Cloud
ImageQwen Image 3.0 Pro
qwen-image-3.0-pro Alibaba's flagship Qwen Image 3.0 model for dense layouts, multilingual text rendering, and photographic detail. Supports text-to-image and image editing up to 2K.
This catalog lists the static presets currently bundled with OctaFuse Gateway for direct Admin import. Gateway still lets you create any custom Provider, Model, or compatible endpoint. Not listed does not mean unsupported.
Alibaba Cloud
Imageqwen-image-3.0-pro Alibaba's flagship Qwen Image 3.0 model for dense layouts, multilingual text rendering, and photographic detail. Supports text-to-image and image editing up to 2K.
Alibaba Cloud
Imageqwen-image-3.0 A faster Qwen Image 3.0 model for general generation and editing. Same 2K ceiling as Pro, with a flat per-image output price.
Alibaba Cloud
Imagewan2.7-image-pro Alibaba Wan 2.7 flagship for text-to-image, sequential sets, multi-reference editing, brand color control, and up to 4K generation.
Alibaba Cloud
Imagewan2.7-image A faster Wan 2.7 model for generation, sequential sets, and multi-reference editing, with a 2K output ceiling.
Alibaba Cloud
LLMqwen3.8-max A 2.4T-parameter MoE flagship for coding, office productivity, vision, and long-horizon agents. Native multimodal understanding supports long documents and video, with strong end-to-end delivery on professional tasks.
Alibaba Cloud
LLMqwen3.8-max-preview A flagship Qwen preview for complex reasoning, coding, vision, and long-horizon agentic workloads. Preview channel for early access; production traffic should prefer qwen3.8-max.
Alibaba Cloud
LLMqwen3.8-flash A fast multimodal Qwen3.8 model with a native 1M context window for coding, agents, documents, and visual understanding. Near-flagship quality at a fraction of Qwen3.8 Max cost, with image and video input.
Alibaba Cloud
LLMqwen3.7-plus A cost-efficient multimodal Qwen model for general reasoning, vision, and production applications. It combines text and image understanding with a cost-conscious design for production-scale multimodal reasoning.
Alibaba Cloud
LLMqwen3.7-max The Qwen3.7 flagship, optimized for coding, productivity, and agent-centric workloads. Its agent-oriented design emphasizes coding, office productivity, and complex multi-step work.
Alibaba Cloud
LLMqwen3.7-flash A fast multimodal Qwen3.7 model with stronger vision, agent execution, and coding throughput than Qwen3.6 Flash. Built for low-latency multimodal understanding and agent workflows.
Alibaba Cloud
LLMqwen3.6-plus A scalable long-context Qwen model balancing strong reasoning with efficient inference. A hybrid of efficient linear attention and sparse expert routing improves scalability and inference efficiency.
Alibaba Cloud
LLMqwen3.6-max A high-capability Qwen model for demanding reasoning, coding, and knowledge work.
Alibaba Cloud
LLMqwen3.6-flash A fast multimodal Qwen model with a large context window for high-throughput workloads. It handles text, image, and video inputs with a long context window for latency-sensitive multimodal applications.
Alibaba Cloud
LLMqwen3.5-plus A balanced Qwen model for general chat, reasoning, coding, and multimodal tasks.
Alibaba Cloud
LLMqwen3.5-flash A low-latency Qwen model for responsive chat, extraction, and lightweight agents.
Alibaba Cloud
LLMqwen3-vl-plus A vision-language model for document, image, screen, and visual reasoning tasks.
Alibaba Cloud
LLMqwen3.5-omni-plus A full-capability Qwen model for understanding text, images, audio, and video.
Alibaba Cloud
LLMqwen3.5-omni-flash A fast omni-modal Qwen model for real-time multimodal understanding and interaction.
Alibaba Cloud
LLMqwen-audio-3.0-asr-flash-streaming DashScope real-time speech recognition for low-latency streaming audio.
Alibaba Cloud
LLMqwen-audio-3.0-asr-flash-filetrans DashScope asynchronous speech recognition for recorded audio files.
Alibaba Cloud
LLMqwen-audio-3.0-asr-flash DashScope synchronous speech recognition for short recorded audio.
Alibaba Cloud
LLMqwen3-asr-flash-realtime DashScope Qwen3 ASR real-time speech recognition over WebSocket sessions.
Alibaba Cloud
LLMqwen3-asr-flash-filetrans DashScope Qwen3 ASR asynchronous speech recognition for recorded audio files.
Alibaba Cloud
LLMqwen3-asr-flash DashScope Qwen3 ASR for synchronous audio transcription.
Alibaba Cloud
LLMfun-asr-realtime DashScope Fun-ASR real-time speech recognition over WebSocket streaming.
Alibaba Cloud
LLMfun-asr DashScope Fun-ASR asynchronous speech recognition for recorded audio files.
Alibaba Cloud
LLMqwen-audio-3.0-tts-plus DashScope high-quality text-to-speech model billed by generated characters.
Alibaba Cloud
LLMqwen-audio-3.0-tts-flash DashScope low-latency text-to-speech model billed by generated characters.
Alibaba Cloud
LLMcosyvoice-v1 DashScope text-to-speech model with streaming synthesis support.
Alibaba Cloud
LLMcosyvoice-v2 DashScope text-to-speech model with streaming synthesis support.
Alibaba Cloud
LLMcosyvoice-v3-flash DashScope low-latency text-to-speech model billed by generated characters.
Alibaba Cloud
LLMcosyvoice-v3-plus DashScope high-quality text-to-speech model billed by generated characters.
Alibaba Cloud
LLMcosyvoice-v3.5-flash DashScope CosyVoice 3.5 Flash TTS with voice cloning/design and instruction control; billed by generated characters.
Alibaba Cloud
LLMcosyvoice-v3.5-plus DashScope CosyVoice 3.5 Plus high-expressiveness TTS with voice cloning/design and instruction control; billed by generated characters.
Anthropic
LLMclaude-fable-5 An Anthropic model built for autonomous knowledge work, coding, and multimodal agents. It is positioned for autonomous knowledge work and coding with multimodal input, reasoning, and tool-oriented workflows.
Anthropic
LLMclaude-sonnet-5 A frontier Sonnet model for coding, professional work, and adaptive agentic reasoning. Adaptive reasoning levels let applications balance frontier coding and agent performance against latency and token use.
Anthropic
LLMclaude-opus-5 Anthropic's flagship for demanding reasoning, coding, visual analysis, and long-running agents. It is especially capable at end-to-end software tasks, code review and bug finding, document and chart analysis, complex office deliverables, tool use, and coordinating parallel subagents.
Anthropic
LLMclaude-opus-4.8 A high-capability Opus model for long-context reasoning, coding, and multimodal analysis. It combines multimodal input, reasoning support, and a long context window for demanding general-purpose workflows.
Anthropic
LLMclaude-opus-4.7 An Opus model designed for long-running asynchronous agents and complex software workflows. The model is designed for asynchronous agents that sustain coding and professional work over extended runs.
Anthropic
LLMclaude-opus-4.6 A powerful Claude model for end-to-end coding and long-running professional workflows. It is optimized for agents that operate across complete software and professional workflows instead of isolated prompts.
Anthropic
LLMclaude-sonnet-4.6 A balanced frontier model for iterative development, agents, and professional knowledge work. Strengths include iterative development, complex codebase navigation, and end-to-end project execution with tools.
Anthropic
LLMclaude-haiku-4.5 A fast, efficient Claude model for latency-sensitive chat, coding, and agent tasks. It targets near-frontier intelligence at substantially lower latency and cost for high-volume interactive workloads.
ByteDance
Imagedoubao-seedream-5-0-pro A premium Seedream model for high-fidelity image generation and professional visual creation.
ByteDance
Imagedoubao-seedream-5-0 A Seedream model for general image generation, editing, and visual design workflows.
ByteDance
LLMdoubao-seed-2-0-mini-260215 A compact Doubao model for fast, cost-sensitive chat and lightweight reasoning.
ByteDance
LLMdoubao-seed-2-0-lite-260215 A lightweight Doubao model optimized for high-throughput text and vision workloads.
ByteDance
LLMdoubao-seed-2-0-pro-260215 A capable Doubao model for advanced reasoning, multimodal understanding, and agents.
ByteDance
LLMdoubao-seed-2-1-pro-260628 A newer Doubao Pro model for complex reasoning, coding, and multimodal production use.
ByteDance
LLMdoubao-seed-2-1-turbo-260628 A low-latency Doubao model for responsive multimodal applications and agents.
ByteDance
LLMdoubao-seed-evolving An evolving Doubao model intended for adaptive reasoning and agentic experimentation.
DeepSeek
LLMdeepseek-v4-pro The official DeepSeek V4 Pro flagship for advanced reasoning, coding, and sustained agent workloads. The 0813 production checkpoint keeps 1M context and dual thinking modes under the same API name.
DeepSeek
LLMdeepseek-v4-flash The official DeepSeek V4 Flash model for fast long-context reasoning, coding, and agents. The 0731 post-training refresh keeps the mixture-of-experts architecture while strengthening agentic coding and tool use.
DeepSeek
LLMdeepseek-v3.2 An efficient DeepSeek model combining strong reasoning with agentic tool use. Its sparse-attention design targets efficient long-context reasoning while preserving strong tool-use performance.
gemini-3.1-flash-image A fast Google model for high-quality image generation and editing with multimodal reasoning. Also known as Nano Banana 2, it combines fast generation with advanced editing and multimodal visual reasoning.
gemini-3-pro-image-preview Google's advanced image generation preview for grounded creation and detailed visual editing. Built on Gemini 3 Pro, it emphasizes grounded creation, stronger multimodal reasoning, and detailed image editing.
gemini-3.7-flash A high-efficiency Gemini workhorse for coding, agents, and knowledge work. It improves first-pass code accuracy, multi-step tool use, and production-ready application output over Gemini 3.6 Flash.
gemini-3.6-flash A high-efficiency Gemini model for coding, agents, and web or application development. It focuses on polished coding and application outputs with fewer unnecessary edits across agentic workflows.
gemini-3.1-pro-preview A frontier Gemini reasoning preview for software engineering and reliable agentic workflows. The preview improves software-engineering performance, agent reliability, and token efficiency across complex workflows.
gemini-3-flash-preview A fast Gemini thinking model for multi-turn chat, coding, and agent workflows. It offers near-Pro reasoning and tool use at Flash-class speed for multi-turn chat, coding, and agents.
gemini-3.5-flash An efficient multimodal Gemini model with strong coding, reasoning, and parallel-agent performance. It brings near-Pro coding and reasoning to a faster tier and is tuned for parallel agent execution.
gemini-3.5-flash-lite A lightweight Gemini model for focused subagents and high-volume production tasks. The lightweight design is especially suited to focused subagents inside larger multi-agent systems.
gemini-3.1-flash-lite A high-efficiency Gemini preview optimized for high-volume, latency-sensitive applications. It is tuned for high-volume workloads while approaching the quality of larger Flash models.
gemini-2.5-pro A capable Gemini reasoning model for coding, mathematics, science, and multimodal analysis. Built-in thinking supports more deliberate answers across advanced coding, mathematics, science, and multimodal analysis.
gemini-2.5-flash A versatile Gemini workhorse balancing reasoning quality, speed, and multimodal support. It balances built-in reasoning with workhorse throughput for coding, mathematics, science, and multimodal production tasks.
gemini-2.5-flash-lite A low-cost Gemini reasoning model optimized for throughput and fast token generation. The model prioritizes ultra-low latency, low cost, high throughput, and fast token generation.
Meituan LongCat
LLMlongcat-2.0 A sparse expert model from Meituan for coding, repository work, and long-horizon agents. Its sparse expert architecture is built for repository-scale coding, long-horizon problem solving, and agentic execution.
MiniMax
LLMminimax-m2.5 A MiniMax model for practical productivity, coding, and real-world digital work. Training across varied digital work environments extends its coding foundation into practical, end-to-end productivity tasks.
MiniMax
LLMminimax-m2.7 An agentic MiniMax model for autonomous productivity and multi-agent collaboration. It emphasizes autonomous productivity, multi-agent collaboration, and continuous improvement in real-world work.
MiniMax
LLMminimax-m3 A long-context multimodal MiniMax model for coding, knowledge work, and agents. The multimodal foundation is designed for long-horizon agents, coding, and knowledge work over extended context.
Moonshot AI
LLMkimi-k2.5 A native multimodal Kimi model for visual coding and coordinated agent workflows. It pairs native multimodal understanding with visual coding and a self-directed agent-swarm approach.
Moonshot AI
LLMkimi-k2.6 A multimodal Kimi model for long-horizon coding, UI generation, and agent orchestration. It targets long-horizon coding, code-driven UI generation, and coordinated multi-agent execution across complex projects.
Moonshot AI
LLMkimi-k2.7-code A coding-focused Kimi model for reliable end-to-end programming across long contexts. The coding-focused design aims to complete end-to-end programming tasks reliably across long contexts.
Moonshot AI
LLMkimi-k3 An open-weight multimodal Kimi model for complex coding, reasoning, and agentic work. The open-weight multimodal reasoning model is positioned for complex coding, knowledge work, and long-running agents.
OpenAI
LLMwhisper-1 OpenAI Whisper (whisper-1) speech-to-text for /v1/audio/transcriptions. Official price $0.006/minute; gateway meters by duration (per second).
OpenAI
LLMgpt-4o-mini-transcribe OpenAI GPT-4o mini speech-to-text for /v1/audio/transcriptions (json/text). Official rates $1.25/$5 per 1M audio tokens; gateway meters by tokens.
OpenAI
LLMgpt-4o-transcribe OpenAI GPT-4o speech-to-text for /v1/audio/transcriptions (json/text). Official rates $2.50/$10 per 1M audio tokens; gateway meters by tokens.
OpenAI
LLMgpt-4o-transcribe-diarize OpenAI GPT-4o transcription with speaker diarization (/v1/audio/transcriptions; json/text/diarized_json). Same token rates as gpt-4o-transcribe ($2.50/$10 per 1M); gateway meters by tokens.
OpenAI
Imagegpt-image-2 OpenAI's high-fidelity image generation and editing model for the dedicated Images API.
OpenAI
LLMgpt-5.2 A frontier GPT model with adaptive reasoning for coding, agents, and long-context work. Adaptive reasoning allocates computation according to task difficulty for stronger agent and long-context performance.
OpenAI
LLMgpt-5.4 A flagship GPT model for advanced reasoning, coding, multimodal input, and professional workflows. It unifies general GPT and coding capabilities for large-context professional work and tool-driven execution.
OpenAI
LLMgpt-5.4-mini A faster GPT-5.4 variant for high-throughput reasoning, coding, and multimodal applications. The smaller variant keeps strong reasoning and coding capabilities while improving speed and throughput.
OpenAI
LLMgpt-5.4-nano A compact GPT model optimized for low-latency, high-volume, and cost-sensitive tasks. It is tuned for speed-critical classification, extraction, routing, and other high-volume lightweight tasks.
OpenAI
LLMgpt-5.5 A frontier GPT model for complex professional work with stronger reasoning and reliability. It strengthens reliability and token efficiency on difficult professional tasks while retaining long-context support.
OpenAI
LLMgpt-5.6 A GPT-5.6 family model for advanced reasoning, coding, and general agent workflows.
OpenAI
LLMgpt-5.6-sol The GPT-5.6 flagship for demanding reasoning, coding, and multi-step agent tasks. The flagship tier is particularly strong at command-line work and complex multi-step coding workflows.
OpenAI
LLMgpt-5.6-terra A balanced GPT-5.6 model for everyday coding, reasoning, and agentic applications. The balanced tier sits between flagship quality and cost efficiency for everyday coding, reasoning, and agents.
OpenAI
LLMgpt-5.6-luna A fast, efficient GPT-5.6 model for chat, classification, and lightweight agents. The efficient tier targets high-volume chat, classification, and lightweight agent workflows with low latency.
StepFun
LLMstep-3.7-flash An efficient multimodal StepFun model for native image and video understanding. Its multimodal mixture-of-experts design combines a large language backbone with native image and video understanding.
Tencent
LLMhy4-preview Tencent's next-generation productivity model with significantly enhanced Agent and complex task execution capabilities. Hy4 preview features 770B total parameters with 49B active parameters and is optimized for agentic, coding, and productivity scenarios, with stronger capabilities in understanding, planning, tool use, and sustained execution for complex tasks.
Tencent
LLMhy3 A Tencent mixture-of-experts model for configurable reasoning and production agents. Configurable reasoning effort lets production systems trade latency for depth across agentic workloads.
Tencent
LLMhy3-preview A high-efficiency Tencent preview model for agent workflows and production evaluation. The preview offers configurable reasoning levels for efficient evaluation and deployment of agent workflows.
xAI (Grok)
Imagegrok-imagine-image-2.0 xAI's Imagine Image 2.0 model for precise generation, editing, and multi-reference creative work. It follows detailed instructions, preserves identity across edits, and supports 1K/2K output.
xAI (Grok)
Imagegrok-imagine-image-quality A high-fidelity xAI model for fast image generation, editing, and reference-guided creation.
xAI (Grok)
LLMgrok-4.6 A frontier Grok model for long-running agents, coding, and knowledge work. It focuses on multi-step agents, ambitious interactive and visual work, and sustained coding trajectories.
xAI (Grok)
LLMgrok-4.5 A general-purpose Grok model for coding, knowledge work, STEM, and demanding reasoning. It supports both reasoning and non-reasoning modes for agentic software and workflow tasks.
xAI (Grok)
LLMgrok-4.3 A multimodal Grok reasoning model for agents, instruction following, and factual tasks. It combines visual input with reasoning for factual instruction following and agent-oriented applications.
Xiaomi Mimo
LLMmimo-v2-omni A Xiaomi omni-modal model for unified text, image, audio, and video understanding.
Xiaomi Mimo
LLMmimo-v2-pro A capable Xiaomi MiMo model for complex reasoning, coding, and agentic workloads.
Xiaomi Mimo
LLMmimo-v2.5-pro Xiaomi's flagship MiMo model for software engineering and long-horizon agent tasks. The flagship emphasizes general agent capability, complex software engineering, and long-horizon execution.
Xiaomi Mimo
LLMmimo-v2.5 A cost-efficient omni-modal MiMo model for agentic and visual understanding workloads. The native omni-modal design aims for Pro-level agent performance with more economical inference and stronger visual understanding.
Zhipu AI
Imageglm-image A Zhipu image model for text-guided generation and general visual creation.
Zhipu AI
LLMglm-5 A flagship open GLM model for systems design, coding, and long-horizon agents. The open foundation model targets complex systems design and production-grade programming across long-running agent workflows.
Zhipu AI
LLMglm-5-turbo A fast GLM model optimized for real-world agent environments and efficient inference. It is deeply optimized for fast inference in real-world, tool-using agent environments.
Zhipu AI
LLMglm-5.1 A GLM model with stronger coding and sustained performance on long-horizon tasks. Its coding improvements focus on sustained, independent execution over tasks that extend far beyond short interactions.
Zhipu AI
LLMglm-5.2 A long-context GLM reasoning model for project-level software engineering and agents. The large-scale reasoning model is intended for project-level software engineering and long-horizon agent workflows.
Zhipu AI
LLMglm-5.3 A flagship GLM model for long-horizon coding, agents, and security review. It keeps the GLM-5.2 base and raises coding and agent performance through post-training.
Zhipu AI
LLMglm-5.3-flash A native multimodal GLM Flash model for visual coding, agents, and professional workflows. It outperforms GLM-5.2 at a fraction of the cost, with native image, video, and file understanding.
Try another keyword or clear the current filters.
Catalog prices come from Gateway presets. Confirm current pricing and availability with the vendor.
Data and contributions
Every model card links to its preset file. Fix the ID, context, modalities, or catalog pricing and open a PR.