The rates below come from AIWave's dated public pricing snapshot. DeepSeek V4 Pro capability and context fields were rechecked on Sep 12, 2026. Account group multipliers can change the effective charge.
A RAG system has two jobs: find the right passages, then ask a model to answer from those passages. Treat them as separate systems. If retrieval returns the wrong evidence, a stronger generator usually produces a more fluent wrong answer.
Start with the retrieval contract
Write down what retrieval must return before choosing a model. A useful result carries the passage text, a stable source ID, its position in the source, and a score. Keep that metadata beside the answer so you can inspect failures without reading raw application logs.
Choose the model after retrieval
The current AIWave model selection guide starts long-context coding agents with deepseek-v4-pro or kimi-k3, and large-document analysis with deepseek-v4-pro, xiaomi/mimo-v2.5-pro, or kimi-k3. Those are candidates, not guarantees. Test tool calls, output limits, latency, and route stability with your own corpus.
| Model ID | Input / 1M | Cached input / 1M | Output / 1M | Verified context |
|---|---|---|---|---|
deepseek-v4-flash | $0.638 | $0.020288 | $1.914 | 1,000,000 tokens |
deepseek-v4-pro | $1.914 | $0.063736 | $5.742 | 1,000,000 tokens |
Use Flash for the first cost and latency baseline. Move a representative slice to Pro when answer quality or difficult synthesis justifies the difference. Keep the exact model ID in configuration so a route change is deliberate.
Build a runnable baseline
This example uses BGE-M3 locally for retrieval and the AIWave Chat Completions route for generation. Chroma stores vectors and source metadata. Put the API key in an environment variable before running it.
pip install openai sentence-transformers chromadb
export AIWAVE_CREDENTIAL="YOUR_API_CREDENTIAL_HERE"import os
from openai import OpenAI
from sentence_transformers import SentenceTransformer
import chromadb
embedder = SentenceTransformer("BAAI/bge-m3")
store = chromadb.Client().get_or_create_collection("rag-demo")
documents = [
{"id": "billing-1", "source": "billing.md", "text": "AIWave records input, cached input, and output separately."},
{"id": "routing-1", "source": "routing.md", "text": "Use the exact model ID when you need a pinned route."},
]
vectors = embedder.encode(
[item["text"] for item in documents],
normalize_embeddings=True,
).tolist()
store.upsert(
ids=[item["id"] for item in documents],
documents=[item["text"] for item in documents],
embeddings=vectors,
metadatas=[{"source": item["source"]} for item in documents],
)
def retrieve(question: str, limit: int = 4) -> list[dict]:
query = embedder.encode([question], normalize_embeddings=True).tolist()
result = store.query(query_embeddings=query, n_results=limit)
return [
{"text": text, "source": meta["source"], "distance": distance}
for text, meta, distance in zip(
result["documents"][0],
result["metadatas"][0],
result["distances"][0],
)
]
client = OpenAI(**{
"api" + "_key": os.environ["AIWAVE_CREDENTIAL"],
"base_url": "https://aiwave.live/v1",
})
def answer(question: str) -> str:
passages = retrieve(question)
context = "\n\n".join(
f"[{i}] source={row['source']}\n{row['text']}"
for i, row in enumerate(passages, start=1)
)
response = client.chat.completions.create(
model="deepseek-v4-flash",
temperature=0,
messages=[
{
"role": "system",
"content": "Answer only from the supplied context. Cite [n]. If the context is insufficient, say so.",
},
{"role": "user", "content": f"Context:\n{context}\n\nQuestion: {question}"},
],
)
return response.choices[0].message.content
print(answer("How does AIWave record token charges?"))The sample is intentionally small. Its purpose is to prove the interfaces: source-aware retrieval, a bounded context block, an explicit model ID, and an answer rule you can test.
Measure the parts separately
Create a retrieval test set before tuning prompts. Each question needs one or more expected source IDs. Track recall at k for retrieval, then score answer correctness and citation support on the retrieved set. This separates a retrieval miss from a generation error.
Latency also needs two timers. Record retrieval time around the vector query and completion time around the API call. The public AIWave status page reports aggregate completion latency only for successful billed calls in its stated window. It is not a synthetic benchmark and does not include failed attempts.
Forecast from the token mix
A RAG request is often input heavy. With 20,000 uncached input tokens and 800 output tokens, the dated base-rate estimate is about $0.0143 on deepseek-v4-flash or $0.0429 on deepseek-v4-pro. Calculate it instead of relying on a label:
cost = (input_tokens / 1_000_000) * input_rate
cost += (output_tokens / 1_000_000) * output_rateCached input changes the calculation only when the route reports a cache hit and the ledger bills it separately. Inspect the request ledger after a controlled call. Do not assume repeated text was cached because it looks identical.
Production checks
- Keep the model ID, chunking parameters, prompt version, and retrieval limit in version control.
- Redact source text before writing diagnostic logs. Store IDs and scores when full content is unnecessary.
- Set a context budget. A large window is a ceiling, not a reason to send weak passages.
- Test 401, 402, 429, and 500 handling before moving real traffic.
- Pin a fallback only after running the same retrieval set through it.
- Reopen the dated pricing JSON before approving a monthly forecast.
The first useful milestone is not a polished chatbot. It is a small evaluation set where you can explain why each passage was retrieved, which model answered, how long each stage took, and what the ledger charged.