Qwen: Qwen3.8 Flash
Chatqwen/qwen3.8-flashQwen3.8 Flash is the lightweight, high-speed natively multimodal model in Alibaba Cloud Bailian's Qwen3.8 series, released on 2026-08-25. It handles image and video understanding and supports deep reasoning (reasoning_content) that can be toggled per request, targeting high-concurrency, low-cost workloads without giving up coding and office tasks. It also supports tool use, prompt caching, and web search. Context window: 1.13M tokens, output: 131K. Being able to switch reasoning off per call makes it practical to run cheap, fast turns and reserve thinking for the requests that need it. Available via OpenAI and Anthropic protocols through Ofox.
Context Window
1M
Max Output Tokens
131K
Released
2026-08-25
Capabilities
VisionFunction CallingReasoningPrompt CachingWeb SearchVideo Input
Available Providers
Aliyunalicloud
Supported Protocols
openaianthropic
Providers
alicloud
Input Tokens
$0.11/M
Output Tokens
$0.39/M
Cache Read
$0.011/M
Cache Write
$0.14/M
Web Search
$0.01/R
Protocols
openai
/v1/chat/completions/v1/responsesanthropic
Aliyun
Input Tokens
$0.11/M
Output Tokens
$0.39/M
Cache Read
$0.011/M
Cache Write
$0.14/M
Web Search
$0.01/R
Protocols
openai
/v1/chat/completions/v1/responsesanthropic
Code Examples
from openai import OpenAIclient = OpenAI(base_url="https://api.ofox.run/v1",api_key="YOUR_OFOX_API_KEY",)response = client.chat.completions.create(model="qwen/qwen3.8-flash",messages=[{"role": "user", "content": "Hello!"}],)print(response.choices[0].message.content)
Uptime & Status
Related Models
Frequently Asked Questions
Qwen: Qwen3.8 Flash on Ofox.ai costs $0.11/M per million input tokens and $0.39/M per million output tokens. Pay-as-you-go, no monthly fees.
Qwen: Qwen3.8 Flash supports a context window of 1M tokens with max output of 131K tokens, allowing you to process large documents and maintain long conversations.
Simply set your base URL to https://api.ofox.run/v1 and use your Ofox API key. The API is OpenAI-compatible — just change the base URL and API key in your existing code.
Qwen: Qwen3.8 Flash supports the following capabilities: Vision, Function Calling, Reasoning, Prompt Caching, Web Search, Video Input. Access all features through the Ofox.ai unified API.