Skip to content

Model Catalog

Compare model capabilities, context, modalities, and catalog pricing.
A starting point, not a limitation

This catalog lists the static presets currently bundled with OctaFuse Gateway for direct Admin import. Gateway still lets you create any custom Provider, Model, or compatible endpoint. Not listed does not mean unsupported.

Learn about custom configuration

74 results

Press / to search

Alibaba Cloud

LLM

Qwen3.8 Max Preview

qwen3.8-max-preview
Context
1M
Max output
64K
Released
2026-07-19

A flagship Qwen preview for complex reasoning, coding, and agentic enterprise workloads.

Modalities
Text Image
Text

Alibaba Cloud

LLM

Qwen3.7 Plus

qwen3.7-plus
Context
1M
Max output
64K
Released
2026-06-02

A cost-efficient multimodal Qwen model for general reasoning, vision, and production applications. It combines text and image understanding with a cost-conscious design for production-scale multimodal reasoning.

Modalities
Text Image Video File
Text

Alibaba Cloud

LLM

Qwen3.7 Max

qwen3.7-max
Context
1M
Max output
64K
Released
2026-05-20

The Qwen3.7 flagship, optimized for coding, productivity, and agent-centric workloads. Its agent-oriented design emphasizes coding, office productivity, and complex multi-step work.

Modalities
Text
Text

Alibaba Cloud

LLM

Qwen3.6 Plus

qwen3.6-plus
Context
1M
Max output
64K
Released
2026-04-02

A scalable long-context Qwen model balancing strong reasoning with efficient inference. A hybrid of efficient linear attention and sparse expert routing improves scalability and inference efficiency.

Modalities
Text Image Video File
Text

Alibaba Cloud

LLM

Qwen3.6 Max

qwen3.6-max
Context
256K
Max output
64K
Released
2026-04-20

A high-capability Qwen model for demanding reasoning, coding, and knowledge work.

Modalities
Text
Text

Alibaba Cloud

LLM

Qwen3.6 flash

qwen3.6-flash
Context
1M
Max output
64K
Released
2026-04-16

A fast multimodal Qwen model with a large context window for high-throughput workloads. It handles text, image, and video inputs with a long context window for latency-sensitive multimodal applications.

Modalities
Text Image Video File
Text

Alibaba Cloud

LLM

Qwen3.5 Plus

qwen3.5-plus
Context
1M
Max output
64K
Released
2026-02-16

A balanced Qwen model for general chat, reasoning, coding, and multimodal tasks.

Modalities
Text Image Video File
Text

Alibaba Cloud

LLM

Qwen3.5 Flash

qwen3.5-flash
Context
1M
Max output
64K
Released
2026-02-16

A low-latency Qwen model for responsive chat, extraction, and lightweight agents.

Modalities
Text Image Video File
Text

Alibaba Cloud

LLM

Qwen3-VL Plus

qwen3-vl-plus
Context
256K
Max output
32K
Released
2025-12-19

A vision-language model for document, image, screen, and visual reasoning tasks.

Modalities
Text Image Video
Text

Alibaba Cloud

LLM

Qwen3.5 Omni Plus

qwen3.5-omni-plus
Context
256K
Max output
64K
Released
2026-03-15

A full-capability Qwen model for understanding text, images, audio, and video.

Modalities
Text Image Audio Video
Text Audio

Alibaba Cloud

LLM

Qwen3.5 Omni Flash

qwen3.5-omni-flash
Context
256K
Max output
64K
Released
2026-03-15

A fast omni-modal Qwen model for real-time multimodal understanding and interaction.

Modalities
Text Image Audio Video
Text Audio

Anthropic

LLM

Claude Fable 5

claude-fable-5
Context
1M
Max output
128K
Released
2026-06-09

An Anthropic model built for autonomous knowledge work, coding, and multimodal agents. It is positioned for autonomous knowledge work and coding with multimodal input, reasoning, and tool-oriented workflows.

Modalities
Text Image File
Text

Anthropic

LLM

Claude Sonnet 5

claude-sonnet-5
Context
1M
Max output
128K
Released
2026-06-30

A frontier Sonnet model for coding, professional work, and adaptive agentic reasoning. Adaptive reasoning levels let applications balance frontier coding and agent performance against latency and token use.

Modalities
Text Image File
Text

Anthropic

LLM

Claude Opus 5

claude-opus-5
Context
1M
Max output
128K
Released
2026-07-24

Anthropic's flagship for demanding reasoning, coding, visual analysis, and long-running agents. It is especially capable at end-to-end software tasks, code review and bug finding, document and chart analysis, complex office deliverables, tool use, and coordinating parallel subagents.

Modalities
Text Image File
Text

Anthropic

LLM

Claude Opus 4.8

claude-opus-4.8
Context
1M
Max output
128K
Released
2026-05-28

A high-capability Opus model for long-context reasoning, coding, and multimodal analysis. It combines multimodal input, reasoning support, and a long context window for demanding general-purpose workflows.

Modalities
Text Image File
Text

Anthropic

LLM

Claude Opus 4.7

claude-opus-4.7
Context
1M
Max output
128K
Released
2026-04-16

An Opus model designed for long-running asynchronous agents and complex software workflows. The model is designed for asynchronous agents that sustain coding and professional work over extended runs.

Modalities
Text Image File
Text

Anthropic

LLM

Claude Opus 4.6

claude-opus-4.6
Context
1M
Max output
128K
Released
2026-02-05

A powerful Claude model for end-to-end coding and long-running professional workflows. It is optimized for agents that operate across complete software and professional workflows instead of isolated prompts.

Modalities
Text Image File
Text

Anthropic

LLM

Claude Sonnet 4.6

claude-sonnet-4.6
Context
1M
Max output
64K
Released
2026-02-17

A balanced frontier model for iterative development, agents, and professional knowledge work. Strengths include iterative development, complex codebase navigation, and end-to-end project execution with tools.

Modalities
Text Image File
Text

Anthropic

LLM

Claude Haiku 4.5

claude-haiku-4.5
Context
200K
Max output
64K
Released
2025-10-15

A fast, efficient Claude model for latency-sensitive chat, coding, and agent tasks. It targets near-frontier intelligence at substantially lower latency and cost for high-volume interactive workloads.

Modalities
Text Image File
Text

ByteDance

Image New

Doubao Seedream 5.0 Pro

doubao-seedream-5-0-pro
Billing
Per image
Released
2026-07-08

A premium Seedream model for high-fidelity image generation and professional visual creation.

Modalities
Text Image
Image

ByteDance

Image

Doubao Seedream 5.0

doubao-seedream-5-0
Billing
Per image
Released
2026-01-28

A Seedream model for general image generation, editing, and visual design workflows.

Modalities
Text Image
Image

ByteDance

LLM

Doubao Seed 2.0 Mini

doubao-seed-2-0-mini-260215
Context
256K
Max output
128K
Released
2026-02-14

A compact Doubao model for fast, cost-sensitive chat and lightweight reasoning.

Modalities
Text Image
Text

ByteDance

LLM

Doubao Seed 2.0 Lite

doubao-seed-2-0-lite-260215
Context
256K
Max output
128K
Released
2026-02-14

A lightweight Doubao model optimized for high-throughput text and vision workloads.

Modalities
Text Image Video Audio
Text

ByteDance

LLM

Doubao Seed 2.0 Pro

doubao-seed-2-0-pro-260215
Context
256K
Max output
128K
Released
2026-02-14

A capable Doubao model for advanced reasoning, multimodal understanding, and agents.

Modalities
Text Image Video
Text

ByteDance

LLM

Doubao Seed 2.1 Pro

doubao-seed-2-1-pro-260628
Context
256K
Max output
256K
Released
2026-06-23

A newer Doubao Pro model for complex reasoning, coding, and multimodal production use.

Modalities
Text Image Video Audio
Text

ByteDance

LLM

Doubao Seed 2.1 Turbo

doubao-seed-2-1-turbo-260628
Context
256K
Max output
256K
Released
2026-06-23

A low-latency Doubao model for responsive multimodal applications and agents.

Modalities
Text Image Video
Text

ByteDance

LLM

Doubao Seed Evolving

doubao-seed-evolving
Context
256K
Max output
256K
Released
2026-06-23

An evolving Doubao model intended for adaptive reasoning and agentic experimentation.

Modalities
Text Image Video
Text

DeepSeek

LLM

DeepSeek-V3.2

deepseek-v3.2
Context
128K
Max output
8K
Released
2025-12-01

An efficient DeepSeek model combining strong reasoning with agentic tool use. Its sparse-attention design targets efficient long-context reasoning while preserving strong tool-use performance.

Modalities
Text
Text

DeepSeek

LLM

DeepSeek V4 Flash

deepseek-v4-flash
Context
1M
Max output
384K
Released
2026-04-24

A fast DeepSeek mixture-of-experts model for long-context reasoning and coding. The efficiency-focused mixture-of-experts architecture activates only a small portion of its parameters for fast long-context inference.

Modalities
Text
Text

DeepSeek

LLM

DeepSeek V4 Pro

deepseek-v4-pro
Context
1M
Max output
384K
Released
2026-04-24

A large DeepSeek mixture-of-experts model for advanced reasoning, coding, and agents. Its large mixture-of-experts design is aimed at advanced reasoning, coding, and sustained agent workloads.

Modalities
Text
Text

Google

Image New

Gemini 3.1 Flash Image

gemini-3.1-flash-image
Billing
Token based
Released
2026-02-01

A fast Google model for high-quality image generation and editing with multimodal reasoning. Also known as Nano Banana 2, it combines fast generation with advanced editing and multimodal visual reasoning.

Modalities
Text Image
Image

Google

Image New

Gemini 3 Pro Image Preview

gemini-3-pro-image-preview
Billing
Token based
Released
2025-11-01

Google's advanced image generation preview for grounded creation and detailed visual editing. Built on Gemini 3 Pro, it emphasizes grounded creation, stronger multimodal reasoning, and detailed image editing.

Modalities
Text Image
Image

Google

LLM

Gemini 3.6 Flash

gemini-3.6-flash
Context
1M
Max output
64K
Released
2026-07-21

A high-efficiency Gemini model for coding, agents, and web or application development. It focuses on polished coding and application outputs with fewer unnecessary edits across agentic workflows.

Modalities
Text Image Audio Video File
Text

Google

LLM

Gemini 3.1 Pro Preview

gemini-3.1-pro-preview
Context
1M
Max output
64K
Released
2026-02-19

A frontier Gemini reasoning preview for software engineering and reliable agentic workflows. The preview improves software-engineering performance, agent reliability, and token efficiency across complex workflows.

Modalities
Text Image Audio Video File
Text

Google

LLM

Gemini 3 Flash Preview

gemini-3-flash-preview
Context
1M
Max output
64K
Released
2025-12-17

A fast Gemini thinking model for multi-turn chat, coding, and agent workflows. It offers near-Pro reasoning and tool use at Flash-class speed for multi-turn chat, coding, and agents.

Modalities
Text Image Audio Video File
Text

Google

LLM

Gemini 3.5 Flash

gemini-3.5-flash
Context
1M
Max output
64K
Released
2026-05-19

An efficient multimodal Gemini model with strong coding, reasoning, and parallel-agent performance. It brings near-Pro coding and reasoning to a faster tier and is tuned for parallel agent execution.

Modalities
Text Image Audio Video File
Text

Google

LLM

Gemini 3.5 Flash Lite

gemini-3.5-flash-lite
Context
1M
Max output
64K
Released
2026-07-21

A lightweight Gemini model for focused subagents and high-volume production tasks. The lightweight design is especially suited to focused subagents inside larger multi-agent systems.

Modalities
Text Image Audio Video File
Text

Google

LLM

Gemini 3.1 Flash Lite Preview

gemini-3.1-flash-lite-preview
Context
1M
Max output
64K
Released
2026-03-03

A high-efficiency Gemini preview optimized for high-volume, latency-sensitive applications. It is tuned for high-volume workloads while approaching the quality of larger Flash models.

Modalities
Text Image Audio Video File
Text

Google

LLM

Gemini 2.5 Pro

gemini-2.5-pro
Context
1M
Max output
64K
Released
2025-06-17

A capable Gemini reasoning model for coding, mathematics, science, and multimodal analysis. Built-in thinking supports more deliberate answers across advanced coding, mathematics, science, and multimodal analysis.

Modalities
Text Image Audio Video File
Text

Google

LLM

Gemini 2.5 Flash

gemini-2.5-flash
Context
1M
Max output
64K
Released
2025-06-17

A versatile Gemini workhorse balancing reasoning quality, speed, and multimodal support. It balances built-in reasoning with workhorse throughput for coding, mathematics, science, and multimodal production tasks.

Modalities
Text Image Audio Video File
Text

Google

LLM

Gemini 2.5 Flash Lite

gemini-2.5-flash-lite
Context
1M
Max output
64K
Released
2025-07-22

A low-cost Gemini reasoning model optimized for throughput and fast token generation. The model prioritizes ultra-low latency, low cost, high throughput, and fast token generation.

Modalities
Text Image Audio Video File
Text

Meituan LongCat

LLM New

LongCat 2.0

longcat-2.0
Context
1M
Max output
128K
Released
2026-06-30

A sparse expert model from Meituan for coding, repository work, and long-horizon agents. Its sparse expert architecture is built for repository-scale coding, long-horizon problem solving, and agentic execution.

Modalities
Text
Text

MiniMax

LLM

MiniMax M2.5

minimax-m2.5
Context
196K
Max output
128K
Released
2026-02-12

A MiniMax model for practical productivity, coding, and real-world digital work. Training across varied digital work environments extends its coding foundation into practical, end-to-end productivity tasks.

Modalities
Text
Text

MiniMax

LLM

MiniMax M2.7

minimax-m2.7
Context
196K
Max output
128K
Released
2026-03-18

An agentic MiniMax model for autonomous productivity and multi-agent collaboration. It emphasizes autonomous productivity, multi-agent collaboration, and continuous improvement in real-world work.

Modalities
Text
Text

MiniMax

LLM

MiniMax M3

minimax-m3
Context
1M
Max output
64K
Released
2026-06-01

A long-context multimodal MiniMax model for coding, knowledge work, and agents. The multimodal foundation is designed for long-horizon agents, coding, and knowledge work over extended context.

Modalities
Text Image Video
Text

Moonshot AI

LLM

Kimi K2.5

kimi-k2.5
Context
256K
Max output
256K
Released
2026-01-27

A native multimodal Kimi model for visual coding and coordinated agent workflows. It pairs native multimodal understanding with visual coding and a self-directed agent-swarm approach.

Modalities
Text Image Video
Text

Moonshot AI

LLM

Kimi K2.6

kimi-k2.6
Context
256K
Max output
256K
Released
2026-04-20

A multimodal Kimi model for long-horizon coding, UI generation, and agent orchestration. It targets long-horizon coding, code-driven UI generation, and coordinated multi-agent execution across complex projects.

Modalities
Text Image Video
Text

Moonshot AI

LLM

Kimi K2.7 Code

kimi-k2.7-code
Context
256K
Max output
256K
Released
2026-06-12

A coding-focused Kimi model for reliable end-to-end programming across long contexts. The coding-focused design aims to complete end-to-end programming tasks reliably across long contexts.

Modalities
Text Image Video
Text

Moonshot AI

LLM

Kimi K3

kimi-k3
Context
1M
Max output
1M
Released
2026-07-16

An open-weight multimodal Kimi model for complex coding, reasoning, and agentic work. The open-weight multimodal reasoning model is positioned for complex coding, knowledge work, and long-running agents.

Modalities
Text Image Video
Text

OpenAI

Image

GPT Image 2

gpt-image-2
Billing
Token based
Released
2026-04-01

OpenAI's high-fidelity image generation and editing model for the dedicated Images API.

Modalities
Text Image
Image

OpenAI

LLM

GPT-5.2

gpt-5.2
Context
400K
Max output
128K
Released
2025-12-11

A frontier GPT model with adaptive reasoning for coding, agents, and long-context work. Adaptive reasoning allocates computation according to task difficulty for stronger agent and long-context performance.

Modalities
Text Image
Text

OpenAI

LLM

GPT-5.4

gpt-5.4
Context
1M
Max output
128K
Released
2026-03-05

A flagship GPT model for advanced reasoning, coding, multimodal input, and professional workflows. It unifies general GPT and coding capabilities for large-context professional work and tool-driven execution.

Modalities
Text Image
Text

OpenAI

LLM

GPT-5.4 mini

gpt-5.4-mini
Context
400K
Max output
128K
Released
2026-03-17

A faster GPT-5.4 variant for high-throughput reasoning, coding, and multimodal applications. The smaller variant keeps strong reasoning and coding capabilities while improving speed and throughput.

Modalities
Text Image
Text

OpenAI

LLM

GPT-5.4 nano

gpt-5.4-nano
Context
400K
Max output
128K
Released
2026-03-17

A compact GPT model optimized for low-latency, high-volume, and cost-sensitive tasks. It is tuned for speed-critical classification, extraction, routing, and other high-volume lightweight tasks.

Modalities
Text Image
Text

OpenAI

LLM

GPT-5.5

gpt-5.5
Context
1M
Max output
128K
Released
2026-04-23

A frontier GPT model for complex professional work with stronger reasoning and reliability. It strengthens reliability and token efficiency on difficult professional tasks while retaining long-context support.

Modalities
Text Image
Text

OpenAI

LLM

GPT-5.6

gpt-5.6
Context
1.05M
Max output
128K
Released
2026-07-09

A GPT-5.6 family model for advanced reasoning, coding, and general agent workflows.

Modalities
Text Image
Text

OpenAI

LLM

GPT-5.6 Sol

gpt-5.6-sol
Context
1.05M
Max output
128K
Released
2026-07-09

The GPT-5.6 flagship for demanding reasoning, coding, and multi-step agent tasks. The flagship tier is particularly strong at command-line work and complex multi-step coding workflows.

Modalities
Text Image
Text

OpenAI

LLM

GPT-5.6 Terra

gpt-5.6-terra
Context
1.05M
Max output
128K
Released
2026-07-09

A balanced GPT-5.6 model for everyday coding, reasoning, and agentic applications. The balanced tier sits between flagship quality and cost efficiency for everyday coding, reasoning, and agents.

Modalities
Text Image
Text

OpenAI

LLM

GPT-5.6 Luna

gpt-5.6-luna
Context
1.05M
Max output
128K
Released
2026-07-09

A fast, efficient GPT-5.6 model for chat, classification, and lightweight agents. The efficient tier targets high-volume chat, classification, and lightweight agent workflows with low latency.

Modalities
Text Image
Text

StepFun

LLM

Step 3.7 Flash

step-3.7-flash
Context
256K
Max output
256K
Released
2026-05-29

An efficient multimodal StepFun model for native image and video understanding. Its multimodal mixture-of-experts design combines a large language backbone with native image and video understanding.

Modalities
Text Image Video
Text

Tencent

LLM

Hy3

hy3
Context
256K
Max output
128K
Released
2026-07-06

A Tencent mixture-of-experts model for configurable reasoning and production agents. Configurable reasoning effort lets production systems trade latency for depth across agentic workloads.

Modalities
Text
Text

Tencent

LLM

Hy3 preview

hy3-preview
Context
256K
Max output
128K
Released
2026-04-23

A high-efficiency Tencent preview model for agent workflows and production evaluation. The preview offers configurable reasoning levels for efficient evaluation and deployment of agent workflows.

Modalities
Text
Text

xAI (Grok)

Image New

Grok Imagine Image Quality

grok-imagine-image-quality
Billing
Per image
Released
2026-05-04

A high-fidelity xAI model for fast image generation, editing, and reference-guided creation.

Modalities
Text Image
Image

xAI (Grok)

LLM New

Grok 4.5

grok-4.5
Context
500K
Max output
128K
Released
2026-07-08

A frontier Grok model for coding, knowledge work, STEM, and demanding reasoning. It is positioned as xAI’s highest-capability option for frontier coding, knowledge work, and STEM tasks.

Modalities
Text Image
Text

xAI (Grok)

LLM

Grok 4.3

grok-4.3
Context
1M
Max output
128K
Released
2026-04-30

A multimodal Grok reasoning model for agents, instruction following, and factual tasks. It combines visual input with reasoning for factual instruction following and agent-oriented applications.

Modalities
Text Image
Text

Xiaomi Mimo

LLM

MiMo V2 Omni

mimo-v2-omni
Context
256K
Max output
128K
Released
2026-03-18

A Xiaomi omni-modal model for unified text, image, audio, and video understanding.

Modalities
Text Image Audio Video File
Text

Xiaomi Mimo

LLM

MiMo V2 Pro

mimo-v2-pro
Context
1.05M
Max output
128K
Released
2026-03-18

A capable Xiaomi MiMo model for complex reasoning, coding, and agentic workloads.

Modalities
Text Image Audio Video File
Text

Xiaomi Mimo

LLM

Mimo V2.5 Pro

mimo-v2.5-pro
Context
1.05M
Max output
128K
Released
2026-04-27

Xiaomi's flagship MiMo model for software engineering and long-horizon agent tasks. The flagship emphasizes general agent capability, complex software engineering, and long-horizon execution.

Modalities
Text Image Audio Video File
Text

Xiaomi Mimo

LLM

MiMo V2.5

mimo-v2.5
Context
1.05M
Max output
128K
Released
2026-04-22

A cost-efficient omni-modal MiMo model for agentic and visual understanding workloads. The native omni-modal design aims for Pro-level agent performance with more economical inference and stronger visual understanding.

Modalities
Text Image Audio Video File
Text

Zhipu AI

Image New

GLM Image

glm-image
Billing
Per image
Released
2026-01-14

A Zhipu image model for text-guided generation and general visual creation.

Modalities
Text Image
Image

Zhipu AI

LLM

GLM-5

glm-5
Context
200K
Max output
128K
Released
2026-02-11

A flagship open GLM model for systems design, coding, and long-horizon agents. The open foundation model targets complex systems design and production-grade programming across long-running agent workflows.

Modalities
Text
Text

Zhipu AI

LLM

GLM-5 Turbo

glm-5-turbo
Context
200K
Max output
128K
Released
2026-03-15

A fast GLM model optimized for real-world agent environments and efficient inference. It is deeply optimized for fast inference in real-world, tool-using agent environments.

Modalities
Text Image
Text

Zhipu AI

LLM

GLM-5.1

glm-5.1
Context
200K
Max output
128K
Released
2026-03-27

A GLM model with stronger coding and sustained performance on long-horizon tasks. Its coding improvements focus on sustained, independent execution over tasks that extend far beyond short interactions.

Modalities
Text
Text

Zhipu AI

LLM

GLM-5.2

glm-5.2
Context
1M
Max output
128K
Released
2026-06-13

A long-context GLM reasoning model for project-level software engineering and agents. The large-scale reasoning model is intended for project-level software engineering and long-horizon agent workflows.

Modalities
Text
Text

Catalog prices come from Gateway presets. Confirm current pricing and availability with the vendor.

Data and contributions

Found incorrect Model data?

Every model card links to its preset file. Fix the ID, context, modalities, or catalog pricing and open a PR.

Contribute on GitHub