myTEA
Every model through one API.
Every token put to work.
Multi-Model Gateway & Intelligent Token Routing Platform
AI Model Gateway & Intelligent Token Router
Every model service is built, operated and guaranteed by myTEA.
# Change base_url only. One key, every model.
curl https://api.mytea.ai/v1/chat/completions \
-H "Authorization: Bearer $MYTEA_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "auto",
"messages": [{"role":"user","content":"..."}]
}'
model: "auto" — leave model selection to myTEA, routed by cost / latency / quality.myTEA is a multi-model gateway and
intelligent token-routing platform for enterprises and developers
Not "one more model endpoint", but a single entry point that is reliable, verifiable and cost-controlled. Every model service is built, operated and guaranteed by myTEA.
Personal AI infrastructure for every developer and enterprise
Your gateway, your policies, your ledger.
The basic unit of metering and scheduling for AI inference
Every model and every node, measured with the same ruler.
Exchange and route across models, nodes and channels
Treat tokens as a standardised compute currency to schedule and settle.
Every kind of model capability, delivered in one place
Self-hosted open-weight flagships + official direct lines to closed-model vendors.
Four real-world pain points of enterprise AI
Too many models, changing too fast
New models arrive every quarter; teams have no bandwidth to evaluate and integrate each one.
Opaque pricing, scattered bills
Multiple vendors, currencies and billing schemes make reconciliation a chore for finance.
Reliability out of your hands
Rate limits, queues and regional outages happen, and the business can only wait.
"Is this the real model?" Nobody can say
Resellers stack on resellers: context gets truncated, caches get swapped, and users have no way to verify.
What enterprises need is not "more model endpoints", but one entry point that is reliable, verifiable and cost-controlled.
Two models in the market: Marketplace vs Self-Operated
A marketplace answers "is it available?". Self-operation answers "is it stable, is it genuine, is it fairly priced?".
| Dimension | OpenRouter (Marketplace) | myTEA (Self-Operated) |
|---|---|---|
| Model | Aggregator: routes to third-party providers | Full-stack self-operated: own compute + self-deployed models + in-house gateway |
| Model source | 60+ external providers onboarded | Own inference clusters + official direct lines to closed-model vendors |
| Compute | Holds no compute | Own GPU clusters (H200 / B300) |
| Quality & SLA | Depends on each provider | Uniform across the platform, end-to-end accountability |
| Pricing logic | Provider price + platform fee | Controllable cost, transparent prepaid tiers |
| Verifiability | Weak (multi-hop forwarding) | Strong: probe verification + call audit |
The marketplace model leaves risk with sellers and buyers. The self-operated model keeps the risk and hands certainty to the customer.
Product overview: three layers, one platform
Access, routing and exchange in one stack, running on myTEA's own inference clusters and official direct API channels to closed-model vendors.
One API, OpenAI-compatible protocol. One key for every model; existing code changes only base_url.
Picks the model by price / speed / quality / context, fails over automatically when a model or node is down, and supports custom enterprise policies.
Model Router · Token RouterUnified top-up, tiered pricing and multi-currency settlement, with usage dashboards, audit and probes included.
Token ExchangeTechnical architecture
Every request goes through the Gateway for invocation and metering. The routing engine picks the model and node based on task profile and enterprise policy.
- Gateway: every request is invoked and metered through the gateway; one key for every model.
- Routing engine: selects the model and node by task profile and enterprise policy; switches over within seconds when a node or model fails.
- Own clusters: models are deployed and operated by myTEA, so version, precision and context length are fully under control.
- Official direct line: closed flagship models are called on the vendor's official API with vendor keys, never substituted or re-forwarded; reachable directly from mainland China.
Why we insist on self-operation
Six kinds of certainty that come from owning everything from the gateway to the GPU. Only then can we be accountable for the result.
End-to-end SLA
From gateway to GPU it is all ours, so failures can be located and commitments can be made.
Controllable pricing
No middlemen stacking margins. Tiered discounts go straight to the customer.
Model consistency
What you call is exactly the version we deployed: never swapped, never downgraded.
Data & compliance
Requests never leave the self-operated path. Data-residency options are available by region or industry.
Verifiable
Probe service plus call logs let customers check context length and caching behaviour themselves.
Deployable on-premises
The same gateway and router can run inside your VPC or data centre.
Three core capabilities
Unified AccessMaaS · One API, Every Model
One API key for every model. OpenAI Chat Completions compatible; existing code changes only base_url.
- Python / Node / Java / Go SDKs, with adapters for LangChain, LlamaIndex and other frameworks.
- Unified model catalogue: names, context lengths, prices, quotas and available regions at a glance.
- Enterprise key management: sub-keys, project quotas, cost allocation by department or application.
- Streaming, function calling, JSON mode and thinking mode are all supported.
Supported
Intelligent Token RoutingModel Router · Token Router
Models are like teas: some energise, some linger, some are cheap and brew many rounds. myTEA blends them for each task.
- Automatic failover on model or node failure, transparent to the customer.
- Priority + fallback chains (primary → fallback1 → fallback2).
- Enterprises can define custom routing policies: pinned models, weighted splits, budget caps, canary experiments.
Routing Dimensions
Token ExchangeUnified Billing · One Ledger
Buy, schedule and settle tokens as a "standardised compute currency". That is the Exchange.
- Top up once, use on every model: the prepaid balance is deducted per token across all models.
- Tiered discounts: tiers step up automatically with monthly usage; contract pricing for key accounts.
- Multi-currency settlement: USD primarily, with CNY and others supported.
- Invoices & contracts: enterprise contracts, monthly billing and full invoicing.
Exchange Console
Self-operated compute and model matrix
Flagship open-weight models on flagship hardware, deployed and operated by us. What you call is the full, unabridged version.
| Own inference cluster (South China) | Scale | Memory per GPU | Compute |
|---|---|---|---|
| NVIDIA H200 | 100 servers · 800 GPUs | 141 GB HBM3e | ≈ 791 PFLOPS |
| NVIDIA B300 | 23 servers · 184 GPUs | 288 GB HBM3e | ≈ 414 PFLOPS |
| Total | 123 servers · 984 GPUs | ≈ 166 TB HBM | ≈ 1.2 EFLOPS |
* Compute assumes 8 GPUs per server at FP16 / BF16 dense throughput. Actual figures follow the delivered configuration.
Deployed model matrix FLAGSHIP MoE
Transparent and verifiable: let customers see for themselves
Everything through the gateway
Every request is invoked and metered at the myTEA Gateway, so bills and logs match one to one.
Probe service
Customers can test context length, cache hits and response characteristics at any time to verify the model is genuine.
Direct-to-vendor commitment
Closed-model traffic is contractually bound to the vendor's official API with vendor keys: no substitution, no re-forwarding.
Call audit
Every call is traceable to model version, node, latency and token count.
SLA commitment
Availability and latency targets are written into the contract, with agreed compensation if they are missed.
Reference pricing
Prepaid, settled primarily in USD. Figures below are illustrative only; final prices and quotas are set by contract.
Per-token pricing + monthly tiers
| Model | Input / Output $ / M tokens | $1K–10K / month | > $10K / month |
|---|---|---|---|
| Kimi K3 | 2.50 / 12.50 | 10% off | 20% off |
| GLM-5.4 | 0.60 / 2.00 | 10% off | 20% off |
| DeepSeek V4 | 0.40 / 0.80 | 10% off | 20% off |
Reference: Kimi K3 vendor / OpenRouter list price $3.00 / $15.00 (as of 2026-07).
Closed models at a discount to list
Claude / OpenAI and other closed models, called on the vendor's official API with vendor keys, at a discount to vendor list price. Reachable directly from mainland China.
Monthly nodes or on-premises deployment
The same gateway and router can be deployed into your VPC or data centre, or reserve dedicated GPU nodes by the month. Quoted on request.
Getting started: three steps to go live
Confirm requirements and sign
Agree on models, monthly volume, channels and direct-access needs, then sign the contract.
Prepay and activate the gateway
Complete prepayment. We activate your myTEA Gateway and keys and configure routing policies.
Go live and verify
Start calling. Enable the probe service and SLA monitoring as needed.
Must every request go through the myTEA Gateway?
Can you guarantee direct access to the vendor?
What does the probe test?
Does using vendor keys conflict with calling every model through one myTEA key?
Do you support on-premises deployment?
Different models have different strengths.
myTEA blends and routes them for every task.
Apply for access or request the full proposal and quotation. We reply within one business day.