Frontier intelligence at a cost you control.
Open-weight inference at up to 90% lower cost for long-horizon agents, batch pipelines, and high-volume tasks. Engineered for efficiency at every level of the stack.





Doubleword drives the 'Efficiency Frontier'
Open-weight models, run Doubleword's most efficient inference stack give a better cost-performance trade-off than any alternative. Allowing our customers to take on more ambitious projects at the lowest possible price.
Intelligence vs. cost
Dotted rings are the cheapest option at each intelligence level.
6 of the 10 models on the efficient frontier run on Doubleword. Hover any point for detail.
| Model | Target use case | Intelligence | Async / 1M out | Batch / 1M out |
|---|---|---|---|---|
| Kimi-K3 | Agentic reasoning, long context | High capability43.8intelligence score 43.8 | $11.25 | $7.50 |
| Qwen3.8-27B | Coding and maths workflows | High efficiency33.9intelligence score 33.9 | $2.25 | $1.50 |
| GLM-5.3-Flash | Multimodal agentic coding | High capability41.9intelligence score 41.9 | $0.38 | $0.25 |
| DeepSeek-V4-Flash-0731 | High-throughput ETL, extraction | High efficiency35intelligence score 35 | $0.14 | $0.09 |
Intelligence via Artificial Analysis (Intelligence Index v4.3). Cost is one task of 1,000 tokens in / 500 out, at 24-hour batch rates for every provider.
Inference for token hungry teams
Scale smarter with latency-based pricing.
The same model at three speeds. Choose realtime, async or 24-hour batch, and pay less the longer you can wait.
Realtime, async and batch inference pricing
Lowest latency, for interactive chat and live sessions.
Switch in minutes.
OpenAI Compatible API
Switch instantly, with zero code rewrites required. Keep your existing OpenAI or Anthropic SDKs, update your base URL to api.doubleword.ai, and instantly run your workloads on our cost-aware open-weight models.
Support from (human) assistants.
Moving trillions of tokens? Our engineers give hands-on help with migration, prompt optimization and queue tuning.
Prove the performance with evals.
Don't guess on quality. Run our pre-built evaluation workflows to rigorously test output accuracy against your current production models before routing live traffic.
Pushing the frontier of Inference Engineering
Doubleword contributes to the frontier of inference optimization research - building every level of the stack from kernels to scheduling to offer the most efficient inference for high volume workloads.
Pricing that scales with your workload.
Choose the model that best fits your business needs.
Self-Service Inference
Whenever you need it, at the speed and cost that best fits your workload.
Get started- Access to all major open-weight models
- Usage-based pricing with prepaid credits
- Prompt caching, tool calling and structured outputs
- Realtime, async and batch inference
- Zero Data Retention (ZDR) included
- Community and standard support
At-Scale Inference
For specialized workloads requiring dedicated engineering and white-glove support.
Talk to us- Everything available to self-service
- Increased rate limits
- Custom models
- Workload-specific inference optimization (e.g., speculators) for improved throughput and reduced token costs
- Region-locked data processing (e.g., US or EU-based inference)
- Engineering support for workload optimization
- Dedicated Slack channel with our engineering team
- White-glove onboarding & Proof of Concept (PoC) support
- Volume pricing
- Signed MSA, DPA, and Custom SLAs
FAQs
Let there be tokens
Try Doubleword's OpenAI compatible APIs and get frontier-level intelligence at a fraction of the cost.








