DeepSeek V4 Pro is the latest flagship model from DeepSeek. Leading performance on coding, reasoning, and multilingual tasks. 128K context window handles entire codebases and long documents.
| Provider | DeepSeek |
|---|---|
| Input Price | $1.09 per 1M tokens |
| Output Price | $2.17 per 1M tokens |
| Context Window | 128K tokens |
| AIWave Endpoint | deepseek-v4-pro |
curl https://aiwave.live/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{"model":"deepseek-v4-pro","messages":[{"role":"user","content":"Hello!"}]}'
These figures come straight from the live rate above — monthly token volume multiplied by the per-million price. Adjust the volumes to match your own traffic.
| Workload | Assumption | Monthly tokens | Monthly cost |
|---|---|---|---|
| Chatbot | ~50k short conversations | 5.0M in / 1.5M out | $8.70 |
| Coding assistant | one developer, full-time | 20.0M in / 4.0M out | $30.45 |
| RAG / document Q&A | long retrieved contexts | 50.0M in / 2.0M out | $58.72 |
Rates are pulled from the live pricing endpoint. Token accounting is returned in the usage field of every non-streaming response.
The endpoint is OpenAI-compatible, so the official SDKs work once you change base_url. Nothing else in your code needs to move.
from openai import OpenAI
client = OpenAI(
api_key="sk-YOUR_KEY",
base_url="https://aiwave.live/v1",
)
response = client.chat.completions.create(
model="deepseek-v4-pro",
messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)
print(response.usage) # exact token counts for cost tracking
Streaming uses the same server-sent-event format as OpenAI:
stream = client.chat.completions.create(
model="deepseek-v4-pro",
messages=[{"role": "user", "content": "Explain vector search"}],
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta.content
if delta:
print(delta, end="", flush=True)
Node.js, LangChain, LlamaIndex, the Vercel AI SDK, Cursor and Cline all accept a custom base URL, so they work with this model without a dedicated provider plugin. Full parameter reference lives in the Chat Completions docs.
Same provider, same API surface — the practical difference is price and capability. Because every model here speaks the OpenAI wire format, switching between them is a one-word change to the model field.
| Model ID | Input / 1M | Output / 1M |
|---|---|---|
deepseek-r1-distill-qwen-14b | $0.0154 | $0.0308 |
deepseek-r1-distill-qwen-32b | $0.0154 | $0.0308 |
deepseek-v3 | $0.154 | $0.308 |
deepseek-v3.2-think | $0.154 | $0.308 |
deepseek-chat | $0.182 | $0.364 |
deepseek-v4-flash | $0.206 | $0.412 |